internetarchive / internetarchive/openlibrary

Data Import Request: Book Farm Publishers (Kolkata) ISBN Dataset

Open
#13,153 0 comments 0 reactions 0 assignees View on GitHub
Lead: @mekarpeles Priority: 3
Dominant language
Python
Stars
6.7k
Forks
2k
Avg merge
2d 19h
Merged PRs (30d)
138

Description

### Description
I have a cleaned dataset of regional Indian books published by **Book Farm** (a prominent publisher based in Kolkata specializing in Bengali literature, graphic novels, and comics). Currently, these books cannot be scanned in reading tracker apps like Bookmory because their ISBNs are missing from global open databases.

I have formatted the data to align perfectly with Open Library's schema (`isbn_13`, `title`, `authors`, `publishers`, `publish_date`, `language`) and converted the language values to standard `ben` codes.

### Dataset Details
- **Publisher:** BOOKFARM (Kolkata, India)
- **Primary Language:** Bengali (`ben`)
- **Format:** CSV (Tab/Comma Delimited)
- **Total Records:** [154 rows]

### Attached File
I have attached the fully formatted CSV file below for the community import bot to process.

[bookfarmkolkatadatabase.csv](https://github.com/user-attachments/files/30013265/bookfarmkolkatadatabase.csv)

Contributor guide

Open the contributing guide

Research direction

Review the attached bookfarmkolkatadatabase.csv and the repository's community import bot workflow first; the issue provides no source file or test entry point. Confirm the 154 records match the stated Open Library fields and ben language codes, then verify that the bot imports them without rejected ISBN records.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
databases
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.