contentlayerdev / contentlayerdev/contentlayer
Feature request: parsing custom content formats
- Dominant language
- TypeScript
- Stars
- 3.5k
- Forks
- 192
- PR merge metrics
- No merged PRs in 30d
Description
Here's a little wild idea which might be within the scope of contentlayer, or not. I was just wondering whether it would be possible to expose certain utilities that would enable users to define their own `defineDocumentType` or `makeSource` pipeline for custom content types.
This would allow users to parse arbitrary content files that might not currently be supported by contentlayer. Some examples:
- markdown derivatives like [rmd](https://rmarkdown.rstudio.com/articles_intro.html), [quarto](https://quarto.org/)
- json derivatives like [juypter notebook](https://nbformat.readthedocs.io/en/latest/)
- protobuf files
- docx, pdf?
For the first 2 types, I think we are currently one step away from including them in the pipeline. We could run a separate program that converts these files to markdown before passing them to the default markdown pipeline e.g. jupyter users could run `jupyter nbconvert --to markdown notebook.ipynb`
Protobuf files are interesting in that they come with their own proto types so the descriptive field properties would mostly be for parsing particular fields of interest.
For arbitrary file types, I guess it would be more convenient to expose certain methods for the user to implement such that it can be defined as a new type and automatically parsed by contentlayer as part of the process.
Contributor guide
Assessment
This issue has not been assessed yet.