bigscience-workshop / bigscience-workshop/data_tooling

Create dataset the_times_of_india

Open
#93 0 comments 0 reactions 0 assignees View on GitHub
data catalog
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Description

- uid: the_times_of_india
- type: primary
- description:
- name: The Times of India
- description:
- homepage: https://timesofindia.indiatimes.com/
- validated: True
- languages:
- language_names:
- English
- language_comments: Indian English
- language_locations:
- Southern Asia
- India
- validated: False
- custodian:
- name: Times Internet Limited (subsidiary of The Times Group)
- in_catalogue:
- type: A commercial entity
- location: India
- contact_name:
- contact_email:
- contact_submitter: False
- additional: https://www.timesinternet.in/
- validated: False
- availability:
- procurement:
- for_download: No - we would need to spontaneously reach out to the current owners/custodians
- download_url:
- download_email: https://www.indiatimes.com/contactus
- licensing:
- has_licenses: Yes
- license_text: Unless otherwise stated, copyright and all intellectual property rights in all material presented on the Site (including but not limited to text, audio, video or graphical images), trademarks and logos appearing on this Site are the property of Times Internet Limited, its parent, affiliates and associates and are protected under applicable Indian laws. You agree not to use any framing techniques to enclose any trademark or logo or other proprietary information of TIL; or remove, conceal or obliterate any copyright or other proprietary notice or any credit-line or date-line on other mark or source identifier included on the Site / Service, including without limitation, the size, color, location or style of all proprietary marks. Any infringement shall be vigorously defended and pursued to the fullest extent permitted by law.
- license_properties:
- copyright - all rights reserved
- license_list:
- unknown: License information unavailable
- pii:
- has_pii: Yes
- generic_pii_likely: very likely
- generic_pii_list:
- names
- dates (birth, death, etc.)
- full-face photographs and comparable images
- numeric_pii_likely: none
- numeric_pii_list:
- sensitive_pii_likely: somewhat likely
- sensitive_pii_list:
- political opinions
- religious or philosophical beliefs
- no_pii_justification_class:
- no_pii_justification_text:
- validated: False
- source_category:
- category_type: website
- category_web: news or magazine website
- category_media:
- validated: False
- media:
- category:
- text
- audiovisual
- image
- text_format:
- .HTML
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type: article
- instance_count: 1M

Contributor guide

No contributing guide indexed for this repository

Research direction

Create the file named the_times_of_india.json using the dataset details provided in the issue body. Check that the file represents the listed source, language, custodian, availability, PII, media, and format metadata, and that the new dataset entry is complete.

Written by the indexing model from the issue text.

Assessment

Tech stack
json
Domain
data
Issue type
Feature
Difficulty
1/5
Estimated time
Under an hour
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.