PHPackages                             xberg-io/xberg - PHPackages - PHPackages  [Skip to content](#main-content)[PHPackages](/)[Directory](/)[Categories](/categories)[Trending](/trending)[Leaderboard](/leaderboard)[Changelog](/changelog)[Analyze](/analyze)[Collections](/collections)[Log in](/login)[Sign up](/register)

1. [Directory](/)
2. /
3. [PDF &amp; Document Generation](/categories/documents)
4. /
5. xberg-io/xberg

ActivePhp-ext[PDF &amp; Document Generation](/categories/documents)

xberg-io/xberg
==============

High-performance document intelligence library

v1.0.14(2w ago)8.9k59↓25%514[10 issues](https://github.com/xberg-io/xberg/issues)[3 PRs](https://github.com/xberg-io/xberg/pulls)MITRustPHP &gt;=8.2CI failing

Since Jun 5Pushed 1mo ago28 watchersCompare

[ Source](https://github.com/xberg-io/xberg)[ Packagist](https://packagist.org/packages/xberg-io/xberg)[ RSS](/packages/xberg-io-xberg/feed)WikiDiscussions main Synced 1w ago

READMEChangelog (10)Dependencies (4)Versions (244)Used By (0)

Xberg
=====

[](#xberg)

 [ ![Bindings](https://camo.githubusercontent.com/23b2c7873e51d39aafa797a5571f07783c505f083f976594e38024d10eb4e433/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f42696e64696e67732d616c65662532302544372539302d303037656336) ](https://github.com/xberg-io/alef) [ ![Rust](https://camo.githubusercontent.com/8d362a183a99d843c4b0def49dab36da2e93a18f3498d12303ef21691147d24d/68747470733a2f2f696d672e736869656c64732e696f2f6372617465732f762f78626572673f6c6162656c3d5275737426636f6c6f723d303037656336) ](https://crates.io/crates/xberg) [ ![Python](https://camo.githubusercontent.com/6ad9ced2f118bd8984c85537a9b0248233a71a62c70df0d510bb7be7f4902733/68747470733a2f2f696d672e736869656c64732e696f2f707970692f762f78626572673f6c6162656c3d507974686f6e26636f6c6f723d303037656336) ](https://pypi.org/project/xberg/) [ ![Node.js](https://camo.githubusercontent.com/e2f062646909f4721074581bbe5168ce67240690f233fa1dd858652a1afb0f25/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f4078626572672d696f2f78626572673f6c6162656c3d4e6f64652e6a7326636f6c6f723d303037656336) ](https://www.npmjs.com/package/@xberg-io/xberg) [ ![WASM](https://camo.githubusercontent.com/028e942f70b4aa78d77dd21b03bb4f7fd9d5a8a9b1f35ef9650a9b927e6874bd/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f4078626572672d696f2f78626572672d7761736d3f6c6162656c3d5741534d26636f6c6f723d303037656336) ](https://www.npmjs.com/package/@xberg-io/xberg-wasm) [ ![Java](https://camo.githubusercontent.com/9a7db1185922b0c56f8160f12dd27637a6c8d88d4aa196027c9d29679c768e0d/68747470733a2f2f696d672e736869656c64732e696f2f6d6176656e2d63656e7472616c2f762f696f2e78626572672f78626572673f6c6162656c3d4a61766126636f6c6f723d303037656336) ](https://central.sonatype.com/artifact/io.xberg/xberg) [ ![Go](https://camo.githubusercontent.com/c8e2de0e624007309c62de3cac6441925bc743c132047300fa27be6db84c4bc9/68747470733a2f2f696d672e736869656c64732e696f2f6769746875622f762f7461672f78626572672d696f2f78626572673f6c6162656c3d476f26636f6c6f723d3030376563362666696c7465723d76312a) ](https://github.com/xberg-io/xberg/tree/main/packages/go) [ ![C#](https://camo.githubusercontent.com/4172e67ab82d45dfa8dba21de59d6b3df1a9d5e578c1fe0ccefc06bd275d6116/68747470733a2f2f696d672e736869656c64732e696f2f6e756765742f762f58626572673f6c6162656c3d4325323326636f6c6f723d303037656336) ](https://www.nuget.org/packages/Xberg/) [ ![PHP](https://camo.githubusercontent.com/ccb264494ba2eaf62af97f689b63f797d9d9d97c4f8821c99213c94e9ebfb1f4/68747470733a2f2f696d672e736869656c64732e696f2f7061636b61676973742f762f78626572672d696f2f78626572673f6c6162656c3d50485026636f6c6f723d303037656336) ](https://packagist.org/packages/xberg-io/xberg) [ ![Ruby](https://camo.githubusercontent.com/45c757cf5f901392ec2e3ffd587f313917ae9ffbbc6d182e8df61e5a47c2d995/68747470733a2f2f696d672e736869656c64732e696f2f67656d2f762f78626572673f6c6162656c3d5275627926636f6c6f723d303037656336) ](https://rubygems.org/gems/xberg) [ ![Elixir](https://camo.githubusercontent.com/6b63365d097d7a2008bf1f651eca7a01389b120786ee9e47a98e9d670a12933f/68747470733a2f2f696d672e736869656c64732e696f2f686578706d2f762f78626572673f6c6162656c3d456c6978697226636f6c6f723d303037656336) ](https://hex.pm/packages/xberg) [ ![Dart](https://camo.githubusercontent.com/db454972df0ed7d47fc77db158cc1223384f1b12efd46408da648146381ba4e8/68747470733a2f2f696d672e736869656c64732e696f2f7075622f762f78626572673f6c6162656c3d4461727426636f6c6f723d303037656336) ](https://pub.dev/packages/xberg) [ ![Kotlin](https://camo.githubusercontent.com/bd408d01cb439e684bca684150bc967d4d8f0d38fcaf8a8b139ab149688f92a4/68747470733a2f2f696d672e736869656c64732e696f2f6d6176656e2d63656e7472616c2f762f696f2e78626572672f78626572672d616e64726f69643f6c6162656c3d4b6f746c696e26636f6c6f723d303037656336) ](https://central.sonatype.com/artifact/io.xberg/xberg-android) [ ![Swift](https://camo.githubusercontent.com/b8e3c9990c66f8b878644a15e0ffd0fba03df291119bca044b94033c83d27246/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f53776966742d53504d2d303037656336) ](https://github.com/xberg-io/xberg/tree/main/packages/swift) [ ![Zig](https://camo.githubusercontent.com/376558d850c9af32b9d23b0211489a0b5c979af703256fa0b7c3e917bd591c53/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f5a69672d7061636b6167652d303037656336) ](https://github.com/xberg-io/xberg/tree/main/packages/zig) [ ![C FFI](https://camo.githubusercontent.com/41b9c58c3810775402965a3a7652da9832ea83391c1ff255081e5551900db8f2/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f432d4646492d303037656336) ](https://github.com/xberg-io/xberg/releases) [ ![Docker](https://camo.githubusercontent.com/7d83f05278efa79f40cf820ee19ebf22bd638c94cb23b09bb354e7d1e3282496/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f636b65722d676863722e696f2d3030376563363f6c6f676f3d646f636b6572266c6f676f436f6c6f723d7768697465) ](https://github.com/xberg-io/xberg/pkgs/container/xberg) [ ![License](https://camo.githubusercontent.com/0cd4d42d83d2124c29737dd1519425c87c4b465016ef0cee20cbcb8ef420c0e0/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4d49542d303037656336) ](https://github.com/xberg-io/xberg/blob/main/LICENSE) [ ![Documentation](https://camo.githubusercontent.com/e711516ac9c044796ccbad3fb82fc0ef828999e71390c76d188d2a36e796ee5f/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f63732d78626572672d303037656336) ](https://docs.xberg.io) [ ![Hugging Face](https://camo.githubusercontent.com/92d2756a41edd1950fb68935e52cec6b84708b20388b2b4af6e27a430144dcf3/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f48756767696e67253230466163652d58626572672d303037656336) ](https://huggingface.co/xberg-io)

 [ ![Join Discord](https://camo.githubusercontent.com/c3d59355bb5f7fc8224936ba897a06b87e5ba6f52bd08dfd2f7f78bf669c550d/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446973636f72642d436861742d3030376563363f6c6f676f3d646973636f7264266c6f676f436f6c6f723d7768697465) ](https://discord.gg/xt9WY3GnKR) [ ![Live Demo](https://camo.githubusercontent.com/f4ded88061eb110a708c51d9476bd2fe28b1434b3353e4dae241902beef23025/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c69766525323044656d6f2d4f70656e2d3030376563363f6c6f676f3d776562617373656d626c79266c6f676f436f6c6f723d7768697465) ](https://docs.xberg.io/demo.html) [ ![GitHub Stars](https://camo.githubusercontent.com/64f1e7f3b671dc88cc9925418f63ae38d07accaf3963c2acefee9a4b0b7be4f4/68747470733a2f2f696d672e736869656c64732e696f2f6769746875622f73746172732f78626572672d696f2f78626572673f7374796c653d736f6369616c) ](https://github.com/xberg-io/xberg/stargazers)

Extract clean text, tables, and structured data from documents and code — no format detection, no OCR setup, no stitched-together libraries. **One engine, 15 language bindings, runs anywhere.**

> **Xberg is the next iteration of [Kreuzberg](https://github.com/kreuzberg-dev/kreuzberg-v4-lts).** Same document-intelligence engine, rebuilt and rebranded under a fresh v1 line.

**Feed documents → get clean text, tables, metadata, transcripts, code intelligence · Run it library, CLI, REST API, or MCP server · No GPU needed · Stream multi-GB files · Cache results.**

Documents · Images · Spreadsheets · Email · Archives · Code · Audio · Video

[![crates.io](https://camo.githubusercontent.com/f028aa8f3e5cd4cea118dfa9cd4bcadb3d56089ef5e0338c130a872b9c380778/68747470733a2f2f696d672e736869656c64732e696f2f6372617465732f762f78626572673f7374796c653d666c61742d737175617265)](https://crates.io/crates/xberg)[![npm](https://camo.githubusercontent.com/846a39df47ebf15c07ae8adf8aa6d290f7cdb73bd5d746faad7e9134216beed5/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f4078626572672d696f2f78626572673f7374796c653d666c61742d737175617265)](https://www.npmjs.com/package/@xberg-io/xberg)[![PyPI](https://camo.githubusercontent.com/0d55579c9503463ab6b8b6e3e3ecd63779e71039c723c8f9a4d57a6a62f79c3c/68747470733a2f2f696d672e736869656c64732e696f2f707970692f762f78626572673f7374796c653d666c61742d737175617265)](https://pypi.org/project/xberg/)[![License: MIT](https://camo.githubusercontent.com/422db9fd40f5831c765cf6530b6750c081b696bd18d904cf89554df98c676277/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d677265656e3f7374796c653d666c61742d737175617265)](LICENSE)

[Quick start](#installation) · [What you get](#what-you-get) · [Capabilities](#capabilities) · [CLI](#cli-reference) · [Docs](https://docs.xberg.io)

---

[![Extracting clean Markdown from a PDF in the CLI](docs/assets/demos/extract.gif)](docs/assets/demos/extract.gif)

*Feed any document—get structured text. Extract, batch, stream, or crawl.*

[See more ↓](#demos)

---

What you get
------------

[](#what-you-get)

Point Xberg at anything — a PDF, a spreadsheet, a scanned image, an audio file, a source tree — and get back clean, structured content you can use right away. One core does the format detection, reading, and extraction, so you don't assemble a pipeline yourself. Call it from Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, or C FFI, and run it as a library, CLI tool, REST API, or MCP server.

What it doesHow**Extract from 96 formats**PDFs, Office, images, HTML, email, archives, scientific publications, and code — intelligent MIME detection, streaming for large files.**6 output formats**Plain text, Markdown, Djot, HTML, JSON tree structure, or Structured (JSON with OCR metadata and bounding boxes).**Code intelligence**Functions, classes, imports, symbols, docstrings from 306 programming languages. Syntax-aware chunking for RAG pipelines.**Crawl &amp; recurse**Follow URLs, extract documents from within documents (nested archives, embedded PDFs). Auto/Document/Crawl modes.**OCR on demand**Tesseract, PaddleOCR, Candle, or VLM backends — fallback chains, extensible via plugins. Confidence scores. Language auto-detection.**Transcription**Whisper ONNX for audio/video tracks (MP3, M4A, WAV, WebM, MP4).**Embeddings &amp; search**Local (ONNX models) or provider-hosted (OpenAI, Anthropic, Google, 143 providers via liter-llm). Reranking.**Structured outputs**LLM-powered extraction — local (Ollama, LM Studio, vLLM) or remote (OpenAI, Anthropic, Google).**Enrichment**NER, redaction, summarization, translation, QR code detection, page classification, keyword extraction (YAKE/RAKE), language detection, layout detection, table extraction, token reduction (TOON).**Batch &amp; parallel**Process 100s of documents in parallel. Per-file timeouts. Configurable batch concurrency (`max_concurrent_extractions`).**Caching**Content-hash cache keys — skip re-extraction when the file and config are unchanged.**Deployment**Library, CLI (12 commands), REST API (`xberg serve`), MCP server (9 tools, 3 prompts, 4 resources), Docker.---

Demos
-----

[](#demos)

[![Xberg CLI: extract, batch, detect, formats, cache, serve, mcp](docs/assets/demos/cli.gif)](docs/assets/demos/cli.gif)

*The CLI: 12 commands for extraction, caching, serving, and MCP.*

[![OCR from a scanned image with confidence scores and bounding boxes](docs/assets/demos/ocr.gif)](docs/assets/demos/ocr.gif)

*OCR with confidence scores and bounding boxes. Switch backends without code changes.*

[![Crawling a website and extracting all linked documents](docs/assets/demos/crawl.gif)](docs/assets/demos/crawl.gif)

*Web crawl: fetch a page, follow links, extract all documents recursively.*

[![MCP server integration with Claude Desktop showing extraction tools and prompts](docs/assets/demos/mcp.gif)](docs/assets/demos/mcp.gif)

*MCP server: AI agents extract documents, detect formats, warm models, manage cache.*

[![REST API: POST a document, get JSON extraction results with streaming support](docs/assets/demos/serve.gif)](docs/assets/demos/serve.gif)

*REST API: stream large files, get JSON or Markdown, one endpoint for all formats.*

---

Installation
------------

[](#installation)

### Language Packages

[](#language-packages)

**Python**```
pip install xberg
```

See [Python README](https://github.com/xberg-io/xberg/tree/main/packages/python) for full documentation.

**Node.js / TypeScript**```
npm install @xberg-io/xberg
```

See [Node.js README](https://github.com/xberg-io/xberg/tree/main/crates/xberg-node) for full documentation.

**Rust**```
cargo add xberg
```

See [Rust README](https://github.com/xberg-io/xberg/tree/main/crates/xberg) for full documentation.

**Go**```
go get github.com/xberg-io/xberg
```

See [Go README](https://github.com/xberg-io/xberg/tree/main/packages/go) for full documentation.

**Java**Available on Maven Central as `io.xberg:xberg`. See [Java README](https://github.com/xberg-io/xberg/tree/main/packages/java) for the dependency snippet.

**C#**```
dotnet add package Xberg
```

See [C# README](https://github.com/xberg-io/xberg/tree/main/packages/csharp) for full documentation.

**Ruby**```
gem install xberg
```

See [Ruby README](https://github.com/xberg-io/xberg/tree/main/packages/ruby) for full documentation.

**PHP**```
composer require xberg-io/xberg
```

See [PHP README](https://github.com/xberg-io/xberg/tree/main/packages/php) for full documentation.

**Elixir**Add `{:xberg, "~> 1.0"}` to your `mix.exs` dependencies. See [Elixir README](https://github.com/xberg-io/xberg/tree/main/packages/elixir) for full documentation.

**WebAssembly**```
npm install @xberg-io/xberg-wasm
```

See [WebAssembly README](https://github.com/xberg-io/xberg/tree/main/crates/xberg-wasm) for full documentation.

**Kotlin (Android)**Available on Maven Central as `io.xberg:xberg-android`. See [Kotlin README](https://github.com/xberg-io/xberg/tree/main/packages/kotlin-android) for the dependency snippet.

**Swift**Add via Swift Package Manager. See [Swift README](https://github.com/xberg-io/xberg/tree/main/packages/swift) for full documentation.

**Dart / Flutter**```
dart pub add xberg
```

See [Dart README](https://github.com/xberg-io/xberg/tree/main/packages/dart) for full documentation.

**Zig**Add via `zig fetch`. See [Zig README](https://github.com/xberg-io/xberg/tree/main/packages/zig) for full documentation.

**C/C++ (FFI)**Build from source as part of this workspace. See [C (FFI) README](https://github.com/xberg-io/xberg/tree/main/crates/xberg-ffi) for full documentation.

### CLI &amp; Deployment

[](#cli--deployment)

**CLI Tool**```
brew install xberg-io/tap/xberg
```

12 commands: `extract`, `batch`, `detect`, `formats`, `version`, `cache` (stats/clear/manifest/warm), `serve`, `mcp`, `api`, `embed`, `chunk`, `completions`.

See [CLI usage guide](https://docs.xberg.io/cli/usage/) for detailed documentation.

**Docker**```
docker pull ghcr.io/xberg-io/xberg:latest
```

Run in API, CLI, or MCP modes. See [Docker guide](https://docs.xberg.io/guides/docker/) for examples.

**REST API Server**```
xberg serve --host 0.0.0.0 --port 8000
```

One POST endpoint handles all formats. Returns JSON or Markdown. Stream large files. See [API server guide](https://docs.xberg.io/guides/api-server/).

**MCP Server**```
xberg mcp --transport stdio
```

9 tools (extract, extract\_batch, detect\_mime\_type, cache\_stats, list\_formats, cache\_clear, get\_version, cache\_manifest, cache\_warm). 3 prompts (extract\_document, extract\_with\_ocr, semantic\_search). 4 resources (formats, models, OCR languages, embedding presets).

Add to Claude Desktop or Cursor:

```
{
  "mcpServers": {
    "xberg": { "command": "xberg", "args": ["mcp"] }
  }
}
```

See [MCP integration guide](https://docs.xberg.io/guides/mcp-integration/).

### AI Coding Assistants

[](#ai-coding-assistants)

Install the Xberg plugin from [`xberg-io/plugins`](https://github.com/xberg-io/plugins). Ships extraction APIs, OCR backends, configuration, and language conventions.

**Claude Code**```
/plugin marketplace add xberg-io/plugins
/plugin install xberg@xberg

```

**Codex CLI**```
/plugins add https://github.com/xberg-io/plugins

```

Search for `xberg` and select **Install Plugin**.

**Cursor**Settings → Plugins → Add from URL → `https://github.com/xberg-io/plugins`, then select **xberg**.

**Gemini CLI**```
gemini extensions install https://github.com/xberg-io/plugins

```

**Factory Droid**```
droid plugin marketplace add https://github.com/xberg-io/plugins
droid plugin install xberg@xberg

```

**GitHub Copilot CLI**```
copilot plugin marketplace add https://github.com/xberg-io/plugins
copilot plugin install xberg@xberg

```

**opencode**Add to `opencode.json`:

```
{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@xberg-io/opencode-xberg"]
}
```

---

Quick Start
-----------

[](#quick-start)

Extract text from a document:

```
use xberg::{extract, ExtractInput, ExtractionConfig};

#[tokio::main]
async fn main() -> xberg::Result {
    let config = ExtractionConfig::default();
    let output = extract(
        ExtractInput::from_uri("document.pdf"),
        &config
    ).await?;

    println!("{}", output.results[0].content);
    Ok(())
}
```

Common use cases — see [Quick start guide](https://docs.xberg.io/getting-started/quickstart/) for language-specific examples, OCR, batch processing, and API configuration.

---

Capabilities
------------

[](#capabilities)

**Full feature list**### Supported File Formats (96)

[](#supported-file-formats-96)

96 file formats across 8 major categories with intelligent format detection and comprehensive metadata extraction.

#### Office Documents

[](#office-documents)

CategoryFormatsCapabilities**Word Processing**`.docx`, `.docm`, `.doc`, `.dotx`, `.dotm`, `.dot`, `.odt`, `.pages`Full text, tables, images, metadata, styles**Spreadsheets**`.xlsx`, `.xlsm`, `.xlsb`, `.xls`, `.xla`, `.xlam`, `.xltm`, `.xltx`, `.xlt`, `.ods`, `.numbers`Sheet data, formulas, cell metadata, charts**Presentations**`.pptx`, `.pptm`, `.ppt`, `.ppsx`, `.potx`, `.potm`, `.pot`, `.key`Slides, speaker notes, images, metadata**PDF**`.pdf`Text, tables, images, metadata, OCR support**eBooks**`.epub`, `.fb2`Chapters, metadata, embedded resources**Database**`.dbf`Table data extraction, field type support**Hangul**`.hwp`, `.hwpx`Korean document format, text extraction#### Images (OCR-Enabled)

[](#images-ocr-enabled)

CategoryFormatsFeatures**Raster**`.png`, `.jpg`, `.jpeg`, `.gif`, `.webp`, `.bmp`, `.tiff`, `.tif`OCR, table detection, EXIF metadata, dimensions, color space**Advanced**`.jp2`, `.jpx`, `.jpm`, `.mj2`, `.jbig2`, `.jb2`, `.pnm`, `.pbm`, `.pgm`, `.ppm`OCR via pure-Rust JPEG2000 decoder, JBIG2 support, table detection**HEIC family**`.heic`, `.heics`, `.heif`, `.avif`, `.avcs`EXIF metadata, optional pixel decoding**Vector**`.svg`DOM parsing, embedded text, graphics metadata#### Audio &amp; Video

[](#audio--video)

CategoryFormatsFeatures**Audio**`.mp3`, `.mpga`, `.m4a`, `.wav`, `.webm`Whisper transcription**Video audio track**`.mp4`, `.mpeg`, `.webm`Audio-track transcription only#### Web &amp; Data

[](#web--data)

CategoryFormatsFeatures**Markup**`.html`, `.htm`, `.xhtml`, `.xml`, `.svg`DOM parsing, metadata (Open Graph, Twitter Card), link extraction**Structured Data**`.json`, `.yaml`, `.yml`, `.toml`, `.csv`, `.tsv`Schema detection, nested structures, validation**Text &amp; Markdown**`.txt`, `.md`, `.markdown`, `.djot`, `.mdx`, `.rst`, `.org`, `.rtf`CommonMark, GFM, Djot, MDX, reStructuredText, Org Mode#### Email &amp; Archives

[](#email--archives)

CategoryFormatsFeatures**Email**`.eml`, `.msg`, `.pst`Headers, body (HTML/plain), attachments, threading**Archives**`.zip`, `.tar`, `.tgz`, `.gz`, `.7z`File listing, nested archives, metadata, recursive extraction#### Academic &amp; Scientific

[](#academic--scientific)

CategoryFormatsFeatures**Citations**`.bib`, `.ris`, `.nbib`, `.enw`Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX**Scientific**`.tex`, `.latex`, `.typ`, `.typst`, `.jats`, `.ipynb`LaTeX, Typst, Jupyter notebooks, PubMed JATS**Publishing**`.fb2`, `.docbook`, `.dbk`, `.docbook4`, `.docbook5`, `.opml`FictionBook, DocBook XML, OPML outlines### Code Intelligence (306 Languages)

[](#code-intelligence-306-languages)

Extract structure from 306 programming languages via tree-sitter:

FeatureDescription**Structure Extraction**Functions, classes, methods, structs, interfaces, enums**Import/Export Analysis**Module dependencies, re-exports, wildcard imports**Symbol Extraction**Variables, constants, type aliases, properties**Docstring Parsing**Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats**Syntax-Aware Chunking**Split code by semantic boundaries for RAG pipelines**Diagnostics**Parse errors with line/column positionsPowered by [tree-sitter-language-pack](https://github.com/xberg-io/tree-sitter-language-pack).

### Output Formats (6)

[](#output-formats-6)

FormatUse caseExample**Plain**Raw text, no markup`"Chapter 1\nIntroduction"`**Markdown**Readable, structured, RAG-friendly`"# Chapter 1\n## Introduction"`**Djot**Modern lightweight markupSimilar to Markdown but stricter**HTML**Styled, browser-ready`Chapter 1`**JSON**Machine-readable tree structureHierarchical sections with heading levels**Structured**OCR metadata, bounding boxesJSON with `elements[]` containing `{text, bbox, confidence}`### Deployment Modes

[](#deployment-modes)

ModeCommandTransportUse case**Library**`xberg::extract()`Async functionsEmbed in your application**CLI**`xberg extract document.pdf`12 commandsScripts, batch jobs, CI/CD**REST API**`xberg serve`HTTP POSTMicroservice, serverless deployment**MCP Server**`xberg mcp`stdio or HTTPClaude, Cursor, IDE agents**Docker**`docker run ghcr.io/xberg-io/xberg`All modesContainer deployment### OCR Backends

[](#ocr-backends)

- **Tesseract** — Native C FFI (Linux/macOS/Windows) and WASM (browser)
- **PaddleOCR** — ONNX Runtime, mobile-optimized models
- **Candle** — Pure Rust, CPU-only, lightweight
- **VLM** — GPT-4 Vision, Claude Vision, Gemini Vision, or 143 providers via liter-llm

Fallback chains. Extensible via plugin system.

### Embeddings

[](#embeddings)

**Local (ONNX Runtime):**

- Preset models: fast, balanced (default), quality, multilingual
- Dimensions: 384, 768, 1024

**Provider-hosted:**

- OpenAI, Anthropic, Google, Hugging Face, Mistral, Cohere, and 143 providers total
- Via [liter-llm](https://github.com/xberg-io/liter-llm) integration

**Reranking:**

- Local ONNX rerankers (cross-encoder models)
- Provider-hosted: Cohere Rerank, others

### Structured LLM Extraction

[](#structured-llm-extraction)

Local engines: Ollama, LM Studio, vLLM

Remote: OpenAI, Anthropic, Google, Mistral, Cohere, and 143 providers via liter-llm

Schema validation. Temperature, top-p, frequency penalty tuning.

### Enrichment

[](#enrichment)

- **NER** — GLiNER or LLM-based entity recognition
- **Redaction** — Mask PII (phone, email, SSN, credit card, addresses)
- **Summarization** — Document and section summaries via LLM
- **Translation** — Multi-language via LLM
- **Page Classification** — Tag document pages (cover, toc, content, etc.)
- **QR Code Detection** — Extract and decode QR codes from images
- **Keyword Extraction** — YAKE or RAKE algorithms
- **Language Detection** — Detect document language
- **Layout Detection** — RT-DETR + TATR models for document structure
- **Table Extraction** — Cell-level structure and content
- **Token Reduction** — TOON wire format (~30–50% fewer tokens than JSON)

---

CLI Reference
-------------

[](#cli-reference)

**All 12 commands**CommandSubcommandsPurpose`extract`—Extract text from a single document (path, URL, or stdin)`batch`—Extract from multiple documents in parallel`detect`—Identify MIME type of a file`formats`—List all 96 supported formats and MIME types`version`—Show Xberg version`cache``stats`, `clear`, `manifest`, `warm`Manage extraction cache and models`serve`—Start REST API server (default: )`mcp`—Start MCP server (stdio or HTTP transport)`api``schema`Output OpenAPI 3.1 specification`embed`—Generate embeddings for text (local or provider-hosted)`chunk`—Split text into chunks (text, markdown, YAML, or semantic)`completions`—Generate shell completion scriptsRun `xberg --help` or `xberg  --help` for detailed options.

---

Documentation
-------------

[](#documentation)

Full guides, API references for every binding, format reference, and configuration docs live at **[xberg.io](https://docs.xberg.io/)**.

- [Getting Started](https://docs.xberg.io/getting-started/)
- [Quick Start](https://docs.xberg.io/getting-started/quickstart/)
- [Guides](https://docs.xberg.io/guides/)
- [API Reference](https://docs.xberg.io/reference/api/)
- [Format Reference](https://docs.xberg.io/reference/formats/)
- [Live Demo](https://docs.xberg.io/demo.html) (browser, WASM)

---

Contributing
------------

[](#contributing)

Contributions are welcome! See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.

Join our [Discord community](https://discord.gg/xt9WY3GnKR) for questions and discussion.

---

Part of Xberg.dev
-----------------

[](#part-of-xbergdev)

Xberg is one of six open-source projects from Kreuzberg, Inc.:

- [Xberg](https://github.com/xberg-io/xberg) — document intelligence: text, tables, metadata from 91+ formats with optional OCR.
- [Xberg Enterprise](https://github.com/xberg-io/xberg-enterprise) — managed extraction API with SDKs, dashboards, and observability.
- [crawlberg](https://github.com/xberg-io/crawlberg) — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
- [html-to-markdown](https://github.com/xberg-io/html-to-markdown) — fast, lossless HTML→Markdown engine.
- [liter-llm](https://github.com/xberg-io/liter-llm) — universal LLM API client with native bindings for 14 languages and 143 providers.
- [tree-sitter-language-pack](https://github.com/xberg-io/tree-sitter-language-pack) — tree-sitter grammars and code-intelligence primitives.
- [alef](https://github.com/xberg-io/alef) — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.

---

License
-------

[](#license)

MIT License (MIT) — see [LICENSE](LICENSE) for details.

###  Health Score

62

—

FairBetter than 99% of packages

Maintenance94

Actively maintained with recent releases

Popularity45

Moderate usage in the ecosystem

Community35

Small or concentrated contributor base

Maturity66

Established project with proven stability

 Bus Factor1

Top contributor holds 92.3% of commits — single point of failure

How is this calculated?**Maintenance (25%)** — Last commit recency, latest release date, and issue-to-star ratio. Uses a 2-year decay window.

**Popularity (30%)** — Total and monthly downloads, GitHub stars, and forks. Logarithmic scaling prevents top-heavy scores.

**Community (15%)** — Contributors, dependents, forks, watchers, and maintainers. Measures real ecosystem engagement.

**Maturity (30%)** — Project age, version count, PHP version support, and release stability.

###  Release Activity

Cadence

Every ~1 days

Total

186

Last Release

14d ago

Major Versions

4.9.8 → v5.0.0-rc.12026-05-25

4.9.9 → v5.0.0-rc.42026-06-06

PHP version history (3 changes)4.0.0-rc.22PHP ^8.2

4.3.2PHP ^8.4

v5.0.0-rc.1PHP &gt;=8.2

### Community

Maintainers

![](https://avatars.githubusercontent.com/u/30733348?v=4)[Na'aman Hirschfeld](/maintainers/Goldziher)[@Goldziher](https://github.com/Goldziher)

---

Top Contributors

[![Goldziher](https://avatars.githubusercontent.com/u/30733348?v=4)](https://github.com/Goldziher "Goldziher (6205 commits)")[![kh3rld](https://avatars.githubusercontent.com/u/171191586?v=4)](https://github.com/kh3rld "kh3rld (181 commits)")[![dependabot[bot]](https://avatars.githubusercontent.com/in/29110?v=4)](https://github.com/dependabot[bot] "dependabot[bot] (133 commits)")[![v-tan](https://avatars.githubusercontent.com/u/22367932?v=4)](https://github.com/v-tan "v-tan (72 commits)")[![pratik-mahalle](https://avatars.githubusercontent.com/u/124587957?v=4)](https://github.com/pratik-mahalle "pratik-mahalle (39 commits)")[![tobocop2](https://avatars.githubusercontent.com/u/5562156?v=4)](https://github.com/tobocop2 "tobocop2 (25 commits)")[![naderalexan](https://avatars.githubusercontent.com/u/3957852?v=4)](https://github.com/naderalexan "naderalexan (16 commits)")[![yiftachashkenazi](https://avatars.githubusercontent.com/u/154422620?v=4)](https://github.com/yiftachashkenazi "yiftachashkenazi (7 commits)")[![ivanova-gif](https://avatars.githubusercontent.com/u/247880403?v=4)](https://github.com/ivanova-gif "ivanova-gif (6 commits)")[![deepsource-autofix[bot]](https://avatars.githubusercontent.com/in/57168?v=4)](https://github.com/deepsource-autofix[bot] "deepsource-autofix[bot] (4 commits)")[![dergachoff](https://avatars.githubusercontent.com/u/9453404?v=4)](https://github.com/dergachoff "dergachoff (4 commits)")[![saiintbrisson](https://avatars.githubusercontent.com/u/29989290?v=4)](https://github.com/saiintbrisson "saiintbrisson (3 commits)")[![ktos](https://avatars.githubusercontent.com/u/1633261?v=4)](https://github.com/ktos "ktos (2 commits)")[![dereuromark](https://avatars.githubusercontent.com/u/39854?v=4)](https://github.com/dereuromark "dereuromark (2 commits)")[![hoesler](https://avatars.githubusercontent.com/u/1052770?v=4)](https://github.com/hoesler "hoesler (2 commits)")[![HuachunSi](https://avatars.githubusercontent.com/u/43306282?v=4)](https://github.com/HuachunSi "HuachunSi (2 commits)")[![ishbir](https://avatars.githubusercontent.com/u/43928?v=4)](https://github.com/ishbir "ishbir (2 commits)")[![Ayush7614](https://avatars.githubusercontent.com/u/67006255?v=4)](https://github.com/Ayush7614 "Ayush7614 (2 commits)")[![sandmor](https://avatars.githubusercontent.com/u/58484439?v=4)](https://github.com/sandmor "sandmor (2 commits)")[![stavroskladis](https://avatars.githubusercontent.com/u/43149615?v=4)](https://github.com/stavroskladis "stavroskladis (2 commits)")

---

Tags

buncsharpdocument-intelligenceelixirffigolangjavametadata-extractionnodepdf-extractionpdfiumphppythonragrubyrusttable-extractiontesseracttext-extractionwasmpdfextractiontextdocumentOCR

###  Code Quality

TestsPHPUnit

Static AnalysisPHPStan

Code StylePHP CS Fixer

Type Coverage Yes

### Embed Badge

![Health badge](/badges/xberg-io-xberg/health.svg)

```
[![Health](https://phpackages.com/badges/xberg-io-xberg/health.svg)](https://phpackages.com/packages/xberg-io-xberg)
```

###  Alternatives

[smalot/pdfparser

Pdf parser library. Can read and extract information from pdf file.

2.7k43.1M317](/packages/smalot-pdfparser)[tecnickcom/tc-lib-pdf

PHP PDF Library

1.9k611.6k22](/packages/tecnickcom-tc-lib-pdf)[kartik-v/yii2-export

A library to export server/db data in various formats (e.g. excel, html, pdf, csv etc.)

1613.3M37](/packages/kartik-v-yii2-export)[shipfastlabs/parsel

An expressive, driver-based PHP API for local and remote document parsers.

3394.4k](/packages/shipfastlabs-parsel)[vaites/php-apache-tika

Apache Tika bindings for PHP: extracts text from documents and images (with OCR), metadata and more...

1171.5M2](/packages/vaites-php-apache-tika)[tecnickcom/tc-lib-pdf-parser

PHP library to parse PDF documents

38185.8k10](/packages/tecnickcom-tc-lib-pdf-parser)

PHPackages © 2026

[Directory](/)[Categories](/categories)[Trending](/trending)[Changelog](/changelog)[Analyze](/analyze)
