Azure / Azure/MachineLearningNotebooks
Creating a file dataset from a single directory in datastore requires azureml-dataset-runtime?
- Vorherrschende Sprache
- Jupyter Notebook
- Sterne
- 4.4k
- Forks
- 2.6k
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
I am trying to create a file dataset from a single directory in datastore. Im following the code block from
https://learn.microsoft.com/en-us/python/api/azureml-core/azureml.data.dataset_factory.filedatasetfactory?view=azure-ml-py#azureml-data-dataset-factory-filedatasetfactory-from-files
Specifically,
```python
from azureml.core import Dataset, Datastore
# create file dataset from a single file in datastore
datastore = Datastore.get(workspace, 'workspaceblobstore')
# create file dataset from a single directory in datastore
file_dataset_2 = Dataset.File.from_files(path=(datastore, 'image/'))
```
However, when I try to replicate these steps for my own Datastore, I encounter an Import Error
`ImportError: Missing required package "azureml-dataset-runtime", which can be installed by running: "c:\Users\\.conda\envs\\python.exe" -m pip install azureml-dataset-runtime --upgrade`
I am on Python 3.11.3 and I tried installing azureml-dataset-runtime but I encounter a dependency clash which requires me to downgrade to Python 3.8.
Furthermore, from the PyPI page
https://pypi.org/project/azureml-dataset-runtime/
It states that azureml-dataset-runtime is "is internal, and is not intended to be used directly."
Is this intended? I am trying to mount my data for a custom ML training job, using the [as_mount](https://learn.microsoft.com/en-us/python/api/azureml-core/azureml.data.filedataset?view=azure-ml-py#azureml-data-filedataset-as-mount) function from the FileDataset Class. Please let me know if there is a better alternative to mounting data, or am I forced to use Python 3.8?
---
#### Document Details
⚠ *Do not edit this section. It is required for learn.microsoft.com ➟ GitHub issue linking.*
* ID: 091afd7e-72ca-a384-01db-4da4d40a6734
* Version Independent ID: 0f3783bf-ab1f-f0d6-08f3-90becae914e8
* Content: [azureml.data.dataset_factory.FileDatasetFactory class - Azure Machine Learning Python](https://learn.microsoft.com/en-us/python/api/azureml-core/azureml.data.dataset_factory.filedatasetfactory?view=azure-ml-py#azureml-data-dataset-factory-filedatasetfactory-from-files)
* Content Source: [AzureML-Docset/stable/docs-ref-autogen/azureml-core/azureml.data.dataset_factory.FileDatasetFactory.yml](https://github.com/MicrosoftDocs/MachineLearning-Python-pr/blob/live/AzureML-Docset/stable/docs-ref-autogen/azureml-core/azureml.data.dataset_factory.FileDatasetFactory.yml)
* Service: **machine-learning**
* Sub-service: **core**
* GitHub Login: @DebFro
* Microsoft Alias: **debfro**
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Rechercherichtung
Beginne mit dem Dokumentationsbeispiel zu FileDatasetFactory.from_files und dem Einstiegspunkt FileDataset.as_mount; prüfe dann, wie sich das Beispiel mit Python 3.11 und der Abhängigkeit azureml-dataset-runtime verhält. Erledigt ist die Aufgabe, wenn der dokumentierte Workflow für ein einzelnes Verzeichnis eine unterstützte Runtime und einen unterstützten Abhängigkeitspfad hat oder die Dokumentation die Einschränkung und einen alternativen Mounting-Ansatz klar erklärt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- documentation, machine-learning
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 25/100