docext
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
- Type: Framework
- Imported from GitHub
- Popularity: 2.1K GitHub stars
- License: Apache-2.0
- Source: https://github.com/NanoNets/docext
- Repository: https://github.com/nanonets/docext
- Tags: document, document-analysis, document-data-extraction, document-information-extraction, extraction, llm-ocr, llms, machine-learning, nlp, ocr, ocr-benchmark, ocr-onpremise
- Updated: 2026-10-01
More frameworks
- open-webui — User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
- langchain — The agent engineering platform.
- awesome-llm-apps — 100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
- graphify — Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Curs…
- ragflow — RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create…
- PaddleOCR — Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PD…
- crawl4ai — Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cl…
- hello-agents — 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程