Easy-to-use and powerful LLM and SLM library with awesome model zoo.
-
Updated
May 23, 2026 - Python
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
Knowhere extracts, parses, and outputs structured chunks ready for AI Agents and RAG.
ContextGem: Effortless LLM extraction from documents
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
Local-first AI-powered document intelligence platform for investigative journalism
INF Tech's open-source MLLMs for SOTA visual-language understanding and advanced document intelligence.
Knwler is a lightweight Python tool that extracts structured knowledge graphs from documents using AI. Feed it a PDF or text file and receive a richly connected network of entities, relationships, and topics — complete with an interactive HTML report and exports ready for your favorite graph analytics platform.
Production-grade multimodal RAG for financial document intelligence. Chart understanding · hybrid retrieval · numeric guardrails · multi-tenancy · full observability.
XLSX parser for LLMs, RAG, LangChain, LangGraph, CrewAI, Claude, MCP — turns Excel (.xlsx) into citation-ready JSON with formulas, charts, dependency graphs, and token-counted chunks. Open-source Python library (MIT).
Open-source, self-hosted OSINT investigation platform: turn documents into a live, investigated entity graph. Autonomous agent, graph analytics (centrality, communities, pathfinding), keyless-first tool belt.
Privacy-first document intelligence engine — parse PDFs, DOCX, PPTX, XLSX & CSV into AI-ready chunks for RAG pipelines. Includes HITL review, 3-layer memory chat, and a production FastAPI server.
An explainable AI system that combines Graph Intelligence, Vector Search, and Retrieval-Augmented Generation (RAG) to deliver grounded answers and transparent reasoning paths. Includes a FastAPI backend, Streamlit UI, FAISS vector index, and an in-memory knowledge graph for hybrid retrieval and recommendations.
BoundaryNet - A Semi-Automatic Layout Annotation Tool
A fully local document intelligence system that allows users to build a persistent private knowledge base from documents and query it using retrieval augmented generation. The system runs entirely offline with local embeddings, vector search, and LLM based answer generation.
Generative AI document analysis on Amazon Bedrock AgentCore: a Strands agent routes pages through 27 vision specialists via an MCP Gateway, then correlates results into a structured, accessibility-oriented document tree. Eight-model registry (Anthropic, Amazon, OpenAI), specialist wizard, scripted CDK deployment, React UI on ECS with Cognito.
AI-powered document intelligence platform for automated analysis, processing, and insights extraction from various document formats.
Self-hosted AI governance, document intelligence, RAG, privacy and auditable decision evidence — with local-first deployment and optional Full Rizzo PII detection.
Fast document classification and OCR detection. Analyzes any file type to determine if OCR is needed, saving time and money on unnecessary processing.
Domain-specializált, RAG-alapú multi-ágens pre-audit rendszer transzferár-dokumentációk (TP) elemzésére. 🏆 7. helyezés a PwC Hungary AI Hackathon 2026-on.
To associate your repository with the document-intelligence topic, visit your repo's landing page and select "manage topics."