PHPackages                             kreuzberg-dev/kreuzberg - PHPackages - PHPackages  [Skip to content](#main-content)[PHPackages](/)[Directory](/)[Categories](/categories)[Trending](/trending)[Leaderboard](/leaderboard)[Changelog](/changelog)[Analyze](/analyze)[Collections](/collections)[Log in](/login)[Sign up](/register)

1. [Directory](/)
2. /
3. [PDF &amp; Document Generation](/categories/documents)
4. /
5. kreuzberg-dev/kreuzberg

ActivePhp-ext[PDF &amp; Document Generation](/categories/documents)

kreuzberg-dev/kreuzberg
=======================

High-performance document intelligence library

4.9.9(2mo ago)8.7k0527[3 issues](https://github.com/kreuzberg-dev/kreuzberg/issues)Elastic-2.0RustPHP ^8.4

Since Dec 29Pushed 1w ago30 watchersCompare

[ Source](https://github.com/kreuzberg-dev/kreuzberg)[ Packagist](https://packagist.org/packages/kreuzberg-dev/kreuzberg)[ Docs](https://kreuzberg.dev)[ RSS](/packages/kreuzberg-dev-kreuzberg/feed)WikiDiscussions main Synced 1mo ago

READMEChangelog (10)Dependencies (3)Versions (134)Used By (0)

Xberg
=====

[](#xberg)

 [ ![Bindings](https://camo.githubusercontent.com/23b2c7873e51d39aafa797a5571f07783c505f083f976594e38024d10eb4e433/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f42696e64696e67732d616c65662532302544372539302d303037656336) ](https://github.com/xberg-io/alef) [ ![Rust](https://camo.githubusercontent.com/8d362a183a99d843c4b0def49dab36da2e93a18f3498d12303ef21691147d24d/68747470733a2f2f696d672e736869656c64732e696f2f6372617465732f762f78626572673f6c6162656c3d5275737426636f6c6f723d303037656336) ](https://crates.io/crates/xberg) [ ![Python](https://camo.githubusercontent.com/6ad9ced2f118bd8984c85537a9b0248233a71a62c70df0d510bb7be7f4902733/68747470733a2f2f696d672e736869656c64732e696f2f707970692f762f78626572673f6c6162656c3d507974686f6e26636f6c6f723d303037656336) ](https://pypi.org/project/xberg/) [ ![Node.js](https://camo.githubusercontent.com/e2f062646909f4721074581bbe5168ce67240690f233fa1dd858652a1afb0f25/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f4078626572672d696f2f78626572673f6c6162656c3d4e6f64652e6a7326636f6c6f723d303037656336) ](https://www.npmjs.com/package/@xberg-io/xberg) [ ![WASM](https://camo.githubusercontent.com/028e942f70b4aa78d77dd21b03bb4f7fd9d5a8a9b1f35ef9650a9b927e6874bd/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f4078626572672d696f2f78626572672d7761736d3f6c6162656c3d5741534d26636f6c6f723d303037656336) ](https://www.npmjs.com/package/@xberg-io/xberg-wasm) [ ![Java](https://camo.githubusercontent.com/9a7db1185922b0c56f8160f12dd27637a6c8d88d4aa196027c9d29679c768e0d/68747470733a2f2f696d672e736869656c64732e696f2f6d6176656e2d63656e7472616c2f762f696f2e78626572672f78626572673f6c6162656c3d4a61766126636f6c6f723d303037656336) ](https://central.sonatype.com/artifact/io.xberg/xberg) [ ![Go](https://camo.githubusercontent.com/c8e2de0e624007309c62de3cac6441925bc743c132047300fa27be6db84c4bc9/68747470733a2f2f696d672e736869656c64732e696f2f6769746875622f762f7461672f78626572672d696f2f78626572673f6c6162656c3d476f26636f6c6f723d3030376563362666696c7465723d76312a) ](https://github.com/xberg-io/xberg/tree/main/packages/go) [ ![C#](https://camo.githubusercontent.com/4172e67ab82d45dfa8dba21de59d6b3df1a9d5e578c1fe0ccefc06bd275d6116/68747470733a2f2f696d672e736869656c64732e696f2f6e756765742f762f58626572673f6c6162656c3d4325323326636f6c6f723d303037656336) ](https://www.nuget.org/packages/Xberg/) [ ![PHP](https://camo.githubusercontent.com/ccb264494ba2eaf62af97f689b63f797d9d9d97c4f8821c99213c94e9ebfb1f4/68747470733a2f2f696d672e736869656c64732e696f2f7061636b61676973742f762f78626572672d696f2f78626572673f6c6162656c3d50485026636f6c6f723d303037656336) ](https://packagist.org/packages/xberg-io/xberg) [ ![Ruby](https://camo.githubusercontent.com/45c757cf5f901392ec2e3ffd587f313917ae9ffbbc6d182e8df61e5a47c2d995/68747470733a2f2f696d672e736869656c64732e696f2f67656d2f762f78626572673f6c6162656c3d5275627926636f6c6f723d303037656336) ](https://rubygems.org/gems/xberg) [ ![Elixir](https://camo.githubusercontent.com/6b63365d097d7a2008bf1f651eca7a01389b120786ee9e47a98e9d670a12933f/68747470733a2f2f696d672e736869656c64732e696f2f686578706d2f762f78626572673f6c6162656c3d456c6978697226636f6c6f723d303037656336) ](https://hex.pm/packages/xberg) [ ![Dart](https://camo.githubusercontent.com/db454972df0ed7d47fc77db158cc1223384f1b12efd46408da648146381ba4e8/68747470733a2f2f696d672e736869656c64732e696f2f7075622f762f78626572673f6c6162656c3d4461727426636f6c6f723d303037656336) ](https://pub.dev/packages/xberg) [ ![Kotlin](https://camo.githubusercontent.com/bd408d01cb439e684bca684150bc967d4d8f0d38fcaf8a8b139ab149688f92a4/68747470733a2f2f696d672e736869656c64732e696f2f6d6176656e2d63656e7472616c2f762f696f2e78626572672f78626572672d616e64726f69643f6c6162656c3d4b6f746c696e26636f6c6f723d303037656336) ](https://central.sonatype.com/artifact/io.xberg/xberg-android) [ ![Swift](https://camo.githubusercontent.com/b8e3c9990c66f8b878644a15e0ffd0fba03df291119bca044b94033c83d27246/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f53776966742d53504d2d303037656336) ](https://github.com/xberg-io/xberg/tree/main/packages/swift) [ ![Zig](https://camo.githubusercontent.com/376558d850c9af32b9d23b0211489a0b5c979af703256fa0b7c3e917bd591c53/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f5a69672d7061636b6167652d303037656336) ](https://github.com/xberg-io/xberg/tree/main/packages/zig) [ ![C FFI](https://camo.githubusercontent.com/41b9c58c3810775402965a3a7652da9832ea83391c1ff255081e5551900db8f2/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f432d4646492d303037656336) ](https://github.com/xberg-io/xberg/releases) [ ![Docker](https://camo.githubusercontent.com/7d83f05278efa79f40cf820ee19ebf22bd638c94cb23b09bb354e7d1e3282496/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f636b65722d676863722e696f2d3030376563363f6c6f676f3d646f636b6572266c6f676f436f6c6f723d7768697465) ](https://github.com/xberg-io/xberg/pkgs/container/xberg) [ ![Helm chart](https://camo.githubusercontent.com/3a017ce6c1c131705b4fe30a0a4799644c83e9a70f9b1efbf5832788d5b177f0/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f48656c6d2d63686172742d3030376563363f6c6f676f3d68656c6d266c6f676f436f6c6f723d7768697465) ](https://docs.xberg.io/guides/kubernetes/) [ ![License](https://camo.githubusercontent.com/0cd4d42d83d2124c29737dd1519425c87c4b465016ef0cee20cbcb8ef420c0e0/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4d49542d303037656336) ](https://github.com/xberg-io/xberg/blob/main/LICENSE) [ ![Documentation](https://camo.githubusercontent.com/e711516ac9c044796ccbad3fb82fc0ef828999e71390c76d188d2a36e796ee5f/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446f63732d78626572672d303037656336) ](https://docs.xberg.io) [ ![Hugging Face](https://camo.githubusercontent.com/92d2756a41edd1950fb68935e52cec6b84708b20388b2b4af6e27a430144dcf3/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f48756767696e67253230466163652d58626572672d303037656336) ](https://huggingface.co/xberg-io)

 [ ![Join Discord](https://camo.githubusercontent.com/c3d59355bb5f7fc8224936ba897a06b87e5ba6f52bd08dfd2f7f78bf669c550d/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f446973636f72642d436861742d3030376563363f6c6f676f3d646973636f7264266c6f676f436f6c6f723d7768697465) ](https://discord.gg/xt9WY3GnKR) [ ![Live Demo](https://camo.githubusercontent.com/f4ded88061eb110a708c51d9476bd2fe28b1434b3353e4dae241902beef23025/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c69766525323044656d6f2d4f70656e2d3030376563363f6c6f676f3d776562617373656d626c79266c6f676f436f6c6f723d7768697465) ](https://docs.xberg.io/demo.html) [ ![GitHub Stars](https://camo.githubusercontent.com/64f1e7f3b671dc88cc9925418f63ae38d07accaf3963c2acefee9a4b0b7be4f4/68747470733a2f2f696d672e736869656c64732e696f2f6769746875622f73746172732f78626572672d696f2f78626572673f7374796c653d736f6369616c) ](https://github.com/xberg-io/xberg/stargazers)

Extract clean text, tables, and structured data from documents and code — no format detection, no OCR setup, no stitched-together libraries. **One engine, 15 language bindings, runs anywhere.**

> **Xberg is the next iteration of [Kreuzberg](https://github.com/kreuzberg-dev/kreuzberg-v4-lts).** Same document-intelligence engine, rebuilt and rebranded under a fresh v1 line.

**Feed documents → get clean text, tables, metadata, transcripts, code intelligence · Run it library, CLI, REST API, or MCP server · No GPU needed · Stream multi-GB files · Cache results.**

Documents · Images · Spreadsheets · Email · Archives · Code · Audio · Video

[![crates.io](https://camo.githubusercontent.com/f028aa8f3e5cd4cea118dfa9cd4bcadb3d56089ef5e0338c130a872b9c380778/68747470733a2f2f696d672e736869656c64732e696f2f6372617465732f762f78626572673f7374796c653d666c61742d737175617265)](https://crates.io/crates/xberg)[![npm](https://camo.githubusercontent.com/846a39df47ebf15c07ae8adf8aa6d290f7cdb73bd5d746faad7e9134216beed5/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f4078626572672d696f2f78626572673f7374796c653d666c61742d737175617265)](https://www.npmjs.com/package/@xberg-io/xberg)[![PyPI](https://camo.githubusercontent.com/0d55579c9503463ab6b8b6e3e3ecd63779e71039c723c8f9a4d57a6a62f79c3c/68747470733a2f2f696d672e736869656c64732e696f2f707970692f762f78626572673f7374796c653d666c61742d737175617265)](https://pypi.org/project/xberg/)[![License: MIT](https://camo.githubusercontent.com/422db9fd40f5831c765cf6530b6750c081b696bd18d904cf89554df98c676277/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d677265656e3f7374796c653d666c61742d737175617265)](LICENSE)

[Quick start](#installation) · [What you get](#what-you-get) · [Capabilities](#capabilities) · [CLI](#cli-reference) · [Docs](https://docs.xberg.io)

---

*Feed any document—get structured text. Extract, batch, stream, or crawl.*

---

What you get
------------

[](#what-you-get)

Point Xberg at anything — a PDF, a spreadsheet, a scanned image, an audio file, a source tree — and get back clean, structured content you can use right away. One core does the format detection, reading, and extraction, so you don't assemble a pipeline yourself. Call it from Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, or C FFI, and run it as a library, CLI tool, REST API, or MCP server.

What it doesHow**Extract from 98 formats**PDFs, Office, images, HTML, email, archives, scientific publications, and code — intelligent MIME detection, streaming for large files.**6 output formats**Plain text, Markdown, Djot, HTML, JSON tree structure, or Structured (JSON with OCR metadata and bounding boxes).**Code intelligence**Functions, classes, imports, symbols, docstrings from 306 programming languages. Syntax-aware chunking for RAG pipelines.**Crawl &amp; recurse**Follow URLs, extract documents from within documents (nested archives, embedded PDFs). Auto/Document/Crawl modes.**OCR on demand**Tesseract, PaddleOCR, Candle, or VLM backends — fallback chains, extensible via plugins. Confidence scores. Language auto-detection.**Transcription**Whisper ONNX for audio/video tracks (MP3, M4A, WAV, WebM, MP4).**Embeddings &amp; search**Local (ONNX models) or provider-hosted (OpenAI, Anthropic, Google, 143 providers via liter-llm). Reranking.**Structured outputs**LLM-powered extraction — local (Ollama, LM Studio, vLLM) or remote (OpenAI, Anthropic, Google).**Enrichment**NER, redaction, summarization, translation, QR code detection, page classification, keyword extraction (YAKE/RAKE), language detection, layout detection, table extraction, token reduction (TOON).**Batch &amp; parallel**Process 100s of documents in parallel. Per-file timeouts. Configurable batch concurrency (`max_concurrent_extractions`).**Caching**Content-hash cache keys — skip re-extraction when the file and config are unchanged.**Deployment**Library, CLI (12 commands), REST API (`xberg serve`), MCP server (9 tools, 3 prompts, 4 resources), Docker.---

Installation
------------

[](#installation)

### Language Packages

[](#language-packages)

**Python**```
pip install xberg
```

See [Python README](https://github.com/xberg-io/xberg/tree/main/packages/python) for full documentation.

**Node.js / TypeScript**```
npm install @xberg-io/xberg
```

See [Node.js README](https://github.com/xberg-io/xberg/tree/main/crates/xberg-node) for full documentation.

**Rust**```
cargo add xberg
```

See [Rust README](https://github.com/xberg-io/xberg/tree/main/crates/xberg) for full documentation.

**Go**```
go get github.com/xberg-io/xberg/packages/go@latest
```

> ⚠️ The repository root is not a Go module — `go get github.com/xberg-io/xberg` will fail. Always target the `/packages/go` subdirectory as shown above.

See [Go README](https://github.com/xberg-io/xberg/tree/main/packages/go) for full documentation.

**Java**Available on Maven Central as `io.xberg:xberg`. See [Java README](https://github.com/xberg-io/xberg/tree/main/packages/java) for the dependency snippet.

**C#**```
dotnet add package Xberg
```

See [C# README](https://github.com/xberg-io/xberg/tree/main/packages/csharp) for full documentation.

**Ruby**```
gem install xberg
```

See [Ruby README](https://github.com/xberg-io/xberg/tree/main/packages/ruby) for full documentation.

**PHP**```
composer require xberg-io/xberg
```

See [PHP README](https://github.com/xberg-io/xberg/tree/main/packages/php) for full documentation.

**Elixir**Add `{:xberg, "~> 1.0"}` to your `mix.exs` dependencies. See [Elixir README](https://github.com/xberg-io/xberg/tree/main/packages/elixir) for full documentation.

**WebAssembly**```
npm install @xberg-io/xberg-wasm
```

See [WebAssembly README](https://github.com/xberg-io/xberg/tree/main/crates/xberg-wasm) for full documentation.

**Kotlin (Android)**Available on Maven Central as `io.xberg:xberg-android`. See [Kotlin README](https://github.com/xberg-io/xberg/tree/main/packages/kotlin-android) for the dependency snippet.

**Swift**Add via Swift Package Manager. See [Swift README](https://github.com/xberg-io/xberg/tree/main/packages/swift) for full documentation.

**Dart / Flutter**```
dart pub add xberg
```

See [Dart README](https://github.com/xberg-io/xberg/tree/main/packages/dart) for full documentation.

**Zig**Add via `zig fetch`. See [Zig README](https://github.com/xberg-io/xberg/tree/main/packages/zig) for full documentation.

**C/C++ (FFI)**Build from source as part of this workspace. See [C (FFI) README](https://github.com/xberg-io/xberg/tree/main/crates/xberg-ffi) for full documentation.

### CLI &amp; Deployment

[](#cli--deployment)

**CLI Tool**```
brew install xberg-io/tap/xberg
```

12 commands: `extract`, `batch`, `detect`, `formats`, `version`, `cache` (stats/clear/manifest/warm), `serve`, `mcp`, `api`, `embed`, `chunk`, `completions`.

See [CLI usage guide](https://docs.xberg.io/cli/usage/) for detailed documentation.

**Docker**```
docker pull ghcr.io/xberg-io/xberg:latest
```

Run in API, CLI, or MCP modes. See [Docker guide](https://docs.xberg.io/guides/docker/) for examples.

**REST API Server**```
xberg serve --host 0.0.0.0 --port 8000
```

One POST endpoint handles all formats. Returns JSON or Markdown. Stream large files. See [API server guide](https://docs.xberg.io/guides/api-server/).

**MCP Server**```
xberg mcp --transport stdio
```

9 tools (extract, extract\_batch, detect\_mime\_type, cache\_stats, list\_formats, cache\_clear, get\_version, cache\_manifest, cache\_warm). 3 prompts (extract\_document, extract\_with\_ocr, semantic\_search). 4 resources (formats, models, OCR languages, embedding presets).

Add to Claude Desktop or Cursor:

```
{
  "mcpServers": {
    "xberg": { "command": "xberg", "args": ["mcp"] }
  }
}
```

See [MCP integration guide](https://docs.xberg.io/guides/mcp-integration/).

### AI Coding Assistants

[](#ai-coding-assistants)

Install the Xberg plugin from [`xberg-io/plugins`](https://github.com/xberg-io/plugins). Ships extraction APIs, OCR backends, configuration, and language conventions.

**Claude Code**```
/plugin marketplace add xberg-io/plugins
/plugin install xberg@xberg

```

**Codex CLI**```
/plugins add https://github.com/xberg-io/plugins

```

Search for `xberg` and select **Install Plugin**.

**Cursor**Settings → Plugins → Add from URL → `https://github.com/xberg-io/plugins`, then select **xberg**.

**Gemini CLI**```
gemini extensions install https://github.com/xberg-io/plugins

```

**Factory Droid**```
droid plugin marketplace add https://github.com/xberg-io/plugins
droid plugin install xberg@xberg

```

**GitHub Copilot CLI**```
copilot plugin marketplace add https://github.com/xberg-io/plugins
copilot plugin install xberg@xberg

```

**opencode**Add to `opencode.json`:

```
{
  "$schema": "https://opencode.ai/config.json",
  "plugin": ["@xberg-io/opencode-xberg"]
}
```

---

Quick Start
-----------

[](#quick-start)

Extract text from a document:

```
use xberg::{extract, ExtractInput, ExtractionConfig};

#[tokio::main]
async fn main() -> xberg::Result {
    let config = ExtractionConfig::default();
    let output = extract(
        ExtractInput::from_uri("document.pdf"),
        &config
    ).await?;

    println!("{}", output.results[0].content);
    Ok(())
}
```

Common use cases — see [Quick start guide](https://docs.xberg.io/getting-started/quickstart/) for language-specific examples, OCR, batch processing, and API configuration.

---

Capabilities
------------

[](#capabilities)

**Full feature list**### Supported File Formats (98)

[](#supported-file-formats-98)

98 file formats across 8 major categories with intelligent format detection and comprehensive metadata extraction.

#### Office Documents

[](#office-documents)

CategoryFormatsCapabilities**Word Processing**`.docx`, `.docm`, `.doc`, `.dotx`, `.dotm`, `.dot`, `.odt`, `.pages`, `.wpd`, `.wp`, `.wp5`, `.wp6`Full text, tables, images, metadata, styles**Spreadsheets**`.xlsx`, `.xlsm`, `.xlsb`, `.xls`, `.xla`, `.xlam`, `.xltm`, `.xltx`, `.xlt`, `.ods`, `.numbers`Sheet data, formulas, cell metadata, charts**Presentations**`.pptx`, `.pptm`, `.ppt`, `.ppsx`, `.potx`, `.potm`, `.pot`, `.odp`, `.key`Slides, speaker notes, images, metadata**PDF**`.pdf`Text, tables, images, metadata, OCR support**eBooks**`.epub`, `.fb2`Chapters, metadata, embedded resources**Database**`.dbf`Table data extraction, field type support**Hangul**`.hwp`, `.hwpx`Korean document format, text extraction#### Images (OCR-Enabled)

[](#images-ocr-enabled)

CategoryFormatsFeatures**Raster**`.png`, `.jpg`, `.jpeg`, `.gif`, `.webp`, `.bmp`, `.tiff`, `.tif`OCR, table detection, EXIF metadata, dimensions, color space**Advanced**`.jp2`, `.jpx`, `.jpm`, `.mj2`, `.jbig2`, `.jb2`, `.pnm`, `.pbm`, `.pgm`, `.ppm`OCR via pure-Rust JPEG2000 decoder, JBIG2 support, table detection**HEIC family**`.heic`, `.heics`, `.heif`, `.avif`, `.avcs`EXIF metadata, optional pixel decoding**Vector**`.svg`DOM parsing, embedded text, graphics metadata#### Audio &amp; Video

[](#audio--video)

CategoryFormatsFeatures**Audio**`.mp3`, `.mpga`, `.m4a`, `.wav`, `.webm`Whisper transcription**Video audio track**`.mp4`, `.mpeg`, `.webm`Audio-track transcription only#### Web &amp; Data

[](#web--data)

CategoryFormatsFeatures**Markup**`.html`, `.htm`, `.xhtml`, `.xml`, `.svg`DOM parsing, metadata (Open Graph, Twitter Card), link extraction**Structured Data**`.json`, `.yaml`, `.yml`, `.toml`, `.csv`, `.tsv`Schema detection, nested structures, validation**Text &amp; Markdown**`.txt`, `.md`, `.markdown`, `.djot`, `.mdx`, `.rst`, `.org`, `.rtf`CommonMark, GFM, Djot, MDX, reStructuredText, Org Mode#### Email &amp; Archives

[](#email--archives)

CategoryFormatsFeatures**Email**`.eml`, `.msg`, `.pst`Headers, body (HTML/plain), attachments, threading**Archives**`.zip`, `.tar`, `.tgz`, `.gz`, `.7z`File listing, nested archives, metadata, recursive extraction#### Academic &amp; Scientific

[](#academic--scientific)

CategoryFormatsFeatures**Citations**`.bib`, `.ris`, `.nbib`, `.enw`Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX**Scientific**`.tex`, `.latex`, `.typ`, `.typst`, `.jats`, `.ipynb`LaTeX, Typst, Jupyter notebooks, PubMed JATS**Publishing**`.fb2`, `.docbook`, `.dbk`, `.docbook4`, `.docbook5`, `.opml`FictionBook, DocBook XML, OPML outlines### Code Intelligence (306 Languages)

[](#code-intelligence-306-languages)

Extract structure from 306 programming languages via tree-sitter:

FeatureDescription**Structure Extraction**Functions, classes, methods, structs, interfaces, enums**Import/Export Analysis**Module dependencies, re-exports, wildcard imports**Symbol Extraction**Variables, constants, type aliases, properties**Docstring Parsing**Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats**Syntax-Aware Chunking**Split code by semantic boundaries for RAG pipelines**Diagnostics**Parse errors with line/column positionsPowered by [tree-sitter-language-pack](https://github.com/xberg-io/tree-sitter-language-pack).

### Output Formats (6)

[](#output-formats-6)

FormatUse caseExample**Plain**Raw text, no markup`"Chapter 1\nIntroduction"`**Markdown**Readable, structured, RAG-friendly`"# Chapter 1\n## Introduction"`**Djot**Modern lightweight markupSimilar to Markdown but stricter**HTML**Styled, browser-ready`Chapter 1`**JSON**Machine-readable tree structureHierarchical sections with heading levels**Structured**OCR metadata, bounding boxesJSON with `elements[]` containing `{text, bbox, confidence}`### Deployment Modes

[](#deployment-modes)

ModeCommandTransportUse case**Library**`xberg::extract()`Async functionsEmbed in your application**CLI**`xberg extract document.pdf`12 commandsScripts, batch jobs, CI/CD**REST API**`xberg serve`HTTP POSTMicroservice, serverless deployment**MCP Server**`xberg mcp`stdio or HTTPClaude, Cursor, IDE agents**Docker**`docker run ghcr.io/xberg-io/xberg`All modesContainer deployment### OCR Backends

[](#ocr-backends)

- **Tesseract** — Native C FFI (Linux/macOS/Windows) and WASM (browser)
- **PaddleOCR** — ONNX Runtime, mobile-optimized models
- **Candle** — Pure Rust, CPU-only, lightweight
- **VLM** — GPT-4 Vision, Claude Vision, Gemini Vision, or 143 providers via liter-llm

Fallback chains. Extensible via plugin system.

### Embeddings

[](#embeddings)

**Local (ONNX Runtime):**

- Preset models: fast, balanced (default), quality, multilingual
- Dimensions: 384, 768, 1024

**Provider-hosted:**

- OpenAI, Anthropic, Google, Hugging Face, Mistral, Cohere, and 143 providers total
- Via [liter-llm](https://github.com/xberg-io/liter-llm) integration

**Reranking:**

- Local ONNX rerankers (cross-encoder models)
- Provider-hosted: Cohere Rerank, others

### Structured LLM Extraction

[](#structured-llm-extraction)

Local engines: Ollama, LM Studio, vLLM

Remote: OpenAI, Anthropic, Google, Mistral, Cohere, and 143 providers via liter-llm

Schema validation. Temperature, top-p, frequency penalty tuning.

### Enrichment

[](#enrichment)

- **NER** — GLiNER or LLM-based entity recognition
- **Redaction** — Mask PII (phone, email, SSN, credit card, addresses)
- **Summarization** — Document and section summaries via LLM
- **Translation** — Multi-language via LLM
- **Page Classification** — Tag document pages (cover, toc, content, etc.)
- **QR Code Detection** — Extract and decode QR codes from images
- **Keyword Extraction** — YAKE or RAKE algorithms
- **Language Detection** — Detect document language
- **Layout Detection** — RT-DETR + TATR models for document structure
- **Table Extraction** — Cell-level structure and content
- **Token Reduction** — TOON wire format (~30–50% fewer tokens than JSON)

---

CLI Reference
-------------

[](#cli-reference)

**All 12 commands**CommandSubcommandsPurpose`extract`—Extract text from a single document (path, URL, or stdin)`batch`—Extract from multiple documents in parallel`detect`—Identify MIME type of a file`formats`—List all 98 supported formats and MIME types`version`—Show Xberg version`cache``stats`, `clear`, `manifest`, `warm`Manage extraction cache and models`serve`—Start REST API server (default: )`mcp`—Start MCP server (stdio or HTTP transport)`api``schema`Output OpenAPI 3.1 specification`embed`—Generate embeddings for text (local or provider-hosted)`chunk`—Split text into chunks (text, markdown, YAML, or semantic)`completions`—Generate shell completion scriptsRun `xberg --help` or `xberg  --help` for detailed options.

---

Documentation
-------------

[](#documentation)

Full guides, API references for every binding, format reference, and configuration docs live at **[xberg.io](https://docs.xberg.io/)**.

- [Getting Started](https://docs.xberg.io/getting-started/installation/)
- [Quick Start](https://docs.xberg.io/getting-started/quickstart/)
- [Guides](https://docs.xberg.io/guides/extraction/)
- [API Reference](https://docs.xberg.io/reference/api-python/)
- [Format Reference](https://docs.xberg.io/reference/formats/)
- [Live Demo](https://docs.xberg.io/demo.html) (browser, WASM)

---

Contributing
------------

[](#contributing)

Contributions are welcome! See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.

Join our [Discord community](https://discord.gg/xt9WY3GnKR) for questions and discussion.

---

Part of Xberg.dev
-----------------

[](#part-of-xbergdev)

Xberg is one of six open-source projects from Kreuzberg, Inc.:

- [Xberg](https://github.com/xberg-io/xberg) — document intelligence: text, tables, metadata from 98+ formats with optional OCR.
- [Xberg Enterprise](https://github.com/xberg-io/xberg-enterprise) — managed extraction API with SDKs, dashboards, and observability.
- [crawlberg](https://github.com/xberg-io/crawlberg) — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
- [html-to-markdown](https://github.com/xberg-io/html-to-markdown) — fast, lossless HTML→Markdown engine.
- [liter-llm](https://github.com/xberg-io/liter-llm) — universal LLM API client with native bindings for 14 languages and 143 providers.
- [tree-sitter-language-pack](https://github.com/xberg-io/tree-sitter-language-pack) — tree-sitter grammars and code-intelligence primitives.
- [alef](https://github.com/xberg-io/alef) — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.

---

License
-------

[](#license)

MIT License (MIT) — see [LICENSE](LICENSE) for details.

###  Health Score

60

—

FairBetter than 98% of packages

Maintenance94

Actively maintained with recent releases

Popularity33

Limited adoption so far

Community35

Small or concentrated contributor base

Maturity70

Established project with proven stability

 Bus Factor1

Top contributor holds 90.5% of commits — single point of failure

How is this calculated?**Maintenance (25%)** — Last commit recency, latest release date, and issue-to-star ratio. Uses a 2-year decay window.

**Popularity (30%)** — Total and monthly downloads, GitHub stars, and forks. Logarithmic scaling prevents top-heavy scores.

**Community (15%)** — Contributors, dependents, forks, watchers, and maintainers. Measures real ecosystem engagement.

**Maturity (30%)** — Project age, version count, PHP version support, and release stability.

###  Release Activity

Cadence

Every ~1 days

Total

123

Last Release

45d ago

Major Versions

4.9.8 → v5.0.0-rc.12026-05-25

4.9.9 → v5.0.0-rc.42026-06-06

PHP version history (3 changes)4.0.0-rc.22PHP ^8.2

4.3.2PHP ^8.4

v5.0.0-rc.1PHP &gt;=8.2

### Community

Maintainers

![](https://avatars.githubusercontent.com/u/30733348?v=4)[Na'aman Hirschfeld](/maintainers/Goldziher)[@Goldziher](https://github.com/Goldziher)

---

Top Contributors

[![Goldziher](https://avatars.githubusercontent.com/u/30733348?v=4)](https://github.com/Goldziher "Goldziher (6903 commits)")[![v-tan](https://avatars.githubusercontent.com/u/22367932?v=4)](https://github.com/v-tan "v-tan (214 commits)")[![kh3rld](https://avatars.githubusercontent.com/u/171191586?v=4)](https://github.com/kh3rld "kh3rld (181 commits)")[![dependabot[bot]](https://avatars.githubusercontent.com/in/29110?v=4)](https://github.com/dependabot[bot] "dependabot[bot] (173 commits)")[![tobocop2](https://avatars.githubusercontent.com/u/5562156?v=4)](https://github.com/tobocop2 "tobocop2 (49 commits)")[![pratik-mahalle](https://avatars.githubusercontent.com/u/124587957?v=4)](https://github.com/pratik-mahalle "pratik-mahalle (39 commits)")[![naderalexan](https://avatars.githubusercontent.com/u/3957852?v=4)](https://github.com/naderalexan "naderalexan (16 commits)")[![yiftachashkenazi](https://avatars.githubusercontent.com/u/154422620?v=4)](https://github.com/yiftachashkenazi "yiftachashkenazi (7 commits)")[![ivanova-gif](https://avatars.githubusercontent.com/u/247880403?v=4)](https://github.com/ivanova-gif "ivanova-gif (6 commits)")[![jamon8888](https://avatars.githubusercontent.com/u/145190?v=4)](https://github.com/jamon8888 "jamon8888 (5 commits)")[![dergachoff](https://avatars.githubusercontent.com/u/9453404?v=4)](https://github.com/dergachoff "dergachoff (4 commits)")[![deepsource-autofix[bot]](https://avatars.githubusercontent.com/in/57168?v=4)](https://github.com/deepsource-autofix[bot] "deepsource-autofix[bot] (4 commits)")[![saiintbrisson](https://avatars.githubusercontent.com/u/29989290?v=4)](https://github.com/saiintbrisson "saiintbrisson (3 commits)")[![Ayush7614](https://avatars.githubusercontent.com/u/67006255?v=4)](https://github.com/Ayush7614 "Ayush7614 (2 commits)")[![dereuromark](https://avatars.githubusercontent.com/u/39854?v=4)](https://github.com/dereuromark "dereuromark (2 commits)")[![hoesler](https://avatars.githubusercontent.com/u/1052770?v=4)](https://github.com/hoesler "hoesler (2 commits)")[![HuachunSi](https://avatars.githubusercontent.com/u/43306282?v=4)](https://github.com/HuachunSi "HuachunSi (2 commits)")[![ishbir](https://avatars.githubusercontent.com/u/43928?v=4)](https://github.com/ishbir "ishbir (2 commits)")[![ktos](https://avatars.githubusercontent.com/u/1633261?v=4)](https://github.com/ktos "ktos (2 commits)")[![MannXo](https://avatars.githubusercontent.com/u/17333793?v=4)](https://github.com/MannXo "MannXo (2 commits)")

---

Tags

buncsharpdocument-intelligenceelixirffigolangjavametadata-extractionnodepdf-extractionpdfiumphppythonragrubyrusttable-extractiontesseracttext-extractionwasmpdfperformancexlsxdocxOCRpptxphp8table-extractionrustTesseracttext extractiondocument-parsingdocument-extractiondocument-intelligencedocument-processingmetadata-extraction

###  Code Quality

TestsPHPUnit

Static AnalysisPHPStan

Code StylePHP CS Fixer

Type Coverage Yes

### Embed Badge

![Health badge](/badges/kreuzberg-dev-kreuzberg/health.svg)

```
[![Health](https://phpackages.com/badges/kreuzberg-dev-kreuzberg/health.svg)](https://phpackages.com/packages/kreuzberg-dev-kreuzberg)
```

###  Alternatives

[gotenberg/gotenberg-php

A PHP client for interacting with Gotenberg, a developer-friendly API for converting numerous document formats into PDF files, and more!

3906.6M32](/packages/gotenberg-gotenberg-php)[vaites/php-apache-tika

Apache Tika bindings for PHP: extracts text from documents and images (with OCR), metadata and more...

1171.5M2](/packages/vaites-php-apache-tika)[paperdoc-dev/paperdoc-lib

A zero-dependency PHP library for generating, parsing and converting documents — PDF, DOCX, XLSX, PPTX, HTML, Markdown, CSV and legacy Office formats

1315.4k](/packages/paperdoc-dev-paperdoc-lib)[mnvx/lowrapper

PHP wrapper over LibreOffice converter

129205.2k](/packages/mnvx-lowrapper)[nilgems/laravel-textract

A Laravel package to extract text from files like DOC, XL, Image, Pdf and more. I've developed this package by inspiring "npm textract".

196.0k](/packages/nilgems-laravel-textract)

PHPackages © 2026

[Directory](/)[Categories](/categories)[Trending](/trending)[Changelog](/changelog)[Analyze](/analyze)
