google-deepmind / google-deepmind/open_x_embodiment

Question about the availability of datasets

Open
#37 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
2k
Forks
127
PR merge metrics
No merged PRs in 30d

Description

Hello, many thanks for this great work jointly done by so many researchers!

I am trying out the RT-X dataset by running the code snippet from the colab tutorial. I successfully installed `tfds-nightly`. My runtime environment (related to tensorflow/tfds)
```
tensorflow-metadata 1.14.0
tfds-nightly 4.9.4.dev202401220044
```

But the dataset of rt-x (e.g. fractal_20220817_data) seems not to be included in the list of supported datasets as the output of the snippet of code said:

```
tensorflow_datasets.core.registered.DatasetNotFoundError: Dataset fractal20220817_data not found.

Check that: - if dataset was added recently, it may only be available in `tfds-nightly` - the dataset name is spelled correctly - dataset class defines all base class abstract methods - the module defining the dataset class is imported

```

Tried-out snippet:

```
import tensorflow_datasets as tfds
import tqdm

# optionally replace the DATASET_NAMES below with the list of filtered datasets from the google sheet
DATASET_NAMES = ['fractal_20220817_data', 'kuka', 'bridge', 'taco_play', 'jaco_play', 'berkeley_cable_routing', 'roboturk', 'nyu_door_opening_surprising_effectiveness', 'viola', 'berkeley_autolab_ur5', 'toto', 'language_table', 'columbia_cairlab_pusht_real', 'stanford_kuka_multimodal_dataset_converted_externally_to_rlds', 'nyu_rot_dataset_converted_externally_to_rlds', 'stanford_hydra_dataset_converted_externally_to_rlds', 'austin_buds_dataset_converted_externally_to_rlds', 'nyu_franka_play_dataset_converted_externally_to_rlds', 'maniskill_dataset_converted_externally_to_rlds', 'furniture_bench_dataset_converted_externally_to_rlds', 'cmu_franka_exploration_dataset_converted_externally_to_rlds', 'ucsd_kitchen_dataset_converted_externally_to_rlds', 'ucsd_pick_and_place_dataset_converted_externally_to_rlds', 'austin_sailor_dataset_converted_externally_to_rlds', 'austin_sirius_dataset_converted_externally_to_rlds', 'bc_z', 'usc_cloth_sim_converted_externally_to_rlds', 'utokyo_pr2_opening_fridge_converted_externally_to_rlds', 'utokyo_pr2_tabletop_manipulation_converted_externally_to_rlds', 'utokyo_saytap_converted_externally_to_rlds', 'utokyo_xarm_pick_and_place_converted_externally_to_rlds', 'utokyo_xarm_bimanual_converted_externally_to_rlds', 'robo_net', 'berkeley_mvp_converted_externally_to_rlds', 'berkeley_rpt_converted_externally_to_rlds', 'kaist_nonprehensile_converted_externally_to_rlds', 'stanford_mask_vit_converted_externally_to_rlds', 'tokyo_u_lsmo_converted_externally_to_rlds', 'dlr_sara_pour_converted_externally_to_rlds', 'dlr_sara_grid_clamp_converted_externally_to_rlds', 'dlr_edan_shared_control_converted_externally_to_rlds', 'asu_table_top_converted_externally_to_rlds', 'stanford_robocook_converted_externally_to_rlds', 'eth_agent_affordances', 'imperialcollege_sawyer_wrist_cam', 'iamlab_cmu_pickup_insert_converted_externally_to_rlds', 'uiuc_d3field', 'utaustin_mutex', 'berkeley_fanuc_manipulation', 'cmu_food_manipulation', 'cmu_play_fusion', 'cmu_stretch', 'berkeley_gnm_recon', 'berkeley_gnm_cory_hall', 'berkeley_gnm_sac_son']
DOWNLOAD_DIR = '~/tensorflow_datasets'

print(f"Downloading {len(DATASET_NAMES)} datasets to {DOWNLOAD_DIR}.")
for dataset_name in tqdm.tqdm(DATASET_NAMES):
_ = tfds.load(dataset_name, data_dir=DOWNLOAD_DIR)
```

The list of available datasets in the command line output:

```
tensorflow_datasets.core.registered.DatasetNotFoundError: Dataset fractal20220817_data not found. Available datasets: - abstract_reasoning - accentdb - aeslc - aflw2k3d - ag_news_subset - ai2_arc - ai2_arc_with_ir - amazon_us_reviews - anli - answer_equivalence - arc - asqa - asset - assin2 - bair_robot_pushing_small - bccd - beans - bee_dataset - beir - big_patent - bigearthnet - billsum - binarized_mnist - binary_alpha_digits - ble_wind_field - blimp - booksum - bool_q - bot_adversarial_dialogue - bucc - caltech101 - caltech_birds2010 - caltech_birds2011 - cardiotox - cars196 - cassava - cats_vs_dogs - celeb_a - celeb_a_hq - chexpert - cifar10 - cifar100 - cifar100_n - cifar10_1 - cifar10_corrupted - cifar10_h - cifar10_n - citrus_leaves - cityscapes - clevr - clic
- cmaterdb - coco - coco_captions - coil100 - colorectal_histology - colorectal_histology_large - common_voice - controlled_noisy_web_labels - corr2cause - crema_d - curated_breast_imaging_ddsm - cycle_gan - databricks_dolly - deep_weeds - dementiabank - diabetic_retinopathy_detection - div2k - dmlab - domainnet - downsampled_imagenet - dsprites - dtd - duke_ultrasound - e2e_cleaned - efron_morris75 - emnist - eurosat - fashion_mnist - flic - food101 - fuss - geirhos_conflict_stimuli - groove - gtzan - gtzan_music_speech - horses_or_humans - i_naturalist2017 - i_naturalist2018 - i_naturalist2021 - imagenet2012 - imagenet2012_corrupted - imagenet2012_fewshot - imagenet2012_multilabel - imagenet2012_real - imagenet2012_subset - imagenet_a - imagenet_lt - imagenet_pi - imagenet_r - imagenet_resized - imagenet_sketch - imagenet_v2 - imagenette - imagewang - imdb_reviews - irc_disentanglement - kddcup99 - kitti - kmnist - lambada - lfw
- librispeech - librispeech_lm - libritts - ljspeech - lm1b - lost_and_found - lsun - lvis - malaria - math_dataset - math_qa - mctaco - media_sum - mlqa - mnist - mnist_corrupted - natural_instructions - natural_questions - natural_questions_open - newsroom - nsynth - nyu_depth_v2 - ogbg_molpcba - omniglot - open_images_challenge2019_detection - open_images_v4 - openbookqa - opinion_abstracts - opinosis - opus - oxford_flowers102 - oxford_iiit_pet - para_crawl - pass - patch_camelyon - paws_wiki - paws_x_wiki - penguins - pet_finder - pg19 - piqa - places365_small - placesfull - plant_leaves - plant_village - plantae_k - protein_net - qa4mre - qasc - quac - quality - quickdraw_bitmap - race - radon - real_toxicity_prompts - reddit - reddit_disentanglement - reddit_tifu - ref_coco
- resisc45 - robomimic_mg - robomimic_mh - robomimic_ph - robonet - rock_paper_scissors - rock_you - s3o4d - salient_span_wikipedia - samsum - savee - scan - scene_parse150 - schema_guided_dialogue - sci_tail - scicite - scientific_papers - segment_anything - sentiment140 - shapes3d - sift1m - simpte - siscore - smallnorb - smartwatch_gestures - snli - so2sat - speech_commands - spoken_digit - squad - stanford_dogs - stanford_online_products - star_cfq - starcraft_video - stl10 - story_cloze - summscreen - sun397 - svhn_cropped - symmetric_solids - tatoeba - ted_hrlr_translate - ted_multi_translate - tedlium - tf_flowers - the300w_lp - tiny_shakespeare - titanic - trec - trivia_qa - uc_merced - universal_dependencies - user_libri_audio - vctk
- visual_domain_decathlon - voc - voxceleb - voxforge - waymo_open_dataset - webvid - wider_face - xtreme_pos - xtreme_s - yes_no
```

Could you please help me to figure out the problem with the availability of the RT-X dataset? Many thanks in advance!

Related issues (potentially):

- https://github.com/google-deepmind/open_x_embodiment/issues/29
- https://github.com/tensorflow/datasets/issues/5203

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.