pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
- Type: Framework
- Imported from GitHub
- Popularity: 1.1K GitHub stars
- License: Apache-2.0
- Source: https://github.com/yfedoseev/pdf_oxide
- Tags: data-extraction, document-processing, fast, image-extraction, llm, markdown, pdf, pdf-editor, pdf-generation, pdf-library, pdf-parser, pdf-to-markdown
- Updated: 2026-10-01
More frameworks
- open-webui — User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
- langchain — The agent engineering platform.
- awesome-llm-apps — 100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
- graphify — Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Curs…
- ragflow — RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create…
- PaddleOCR — Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PD…
- crawl4ai — Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cl…
- hello-agents — 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程