#document-processing (9 Repositories)
Ranked open-source repositories tagged with #document-processing, scored by pull request acceptance likelihood and maintainer engagement velocity.
86.9%
27.9h
9 repositories tagged #document-processing
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
ShapeCrawler/ShapeCrawler
PowerPoint .NET library for reading, modifying, and generating PPTX presentations without Microsoft Office
SylphxAI/pdf-reader-mcp
Give your AI agent eyes for PDFs — structured text, tables, OCR, visual evidence, and page-level citations via MCP. Native Rust, local-first.
formkiq/formkiq-core
Open-source document management platform leveraging AWS managed services. RESTful API for document storage, processing, full-text search, and metadata management. Multi-tenant serverless architecture with auto-scaling... deployed entirely in your AWS account.
docling-project/docling-graph
Transform unstructured documents into validated, rich and queryable knowledge graphs.
run-llama/liteparse
A fast, helpful, and open-source document parser
virgiliojr94/book-to-skill
Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
appautomaton/document-SKILLs
Claude Code and Codex SKILLs for PDF, Excel, Word, and PowerPoint manipulation — extraction, forms, formulas, tracked changes, adapted from Anthropic skills.