Nine open-source tools pull text and structure out of PDFs and images, and the scanned page is what splits them. PaddleOCR, Chandra and Falcon-Perception read it on your own machine with a downloaded model, tables and formulas included. Of the three, only PaddleOCR documents running on a CPU, and Chandra's weights are free for commercial use only below 2 million dollars in revenue or funding. Tesseract.js reads printed text from an image, in the browser or in Node.js, with no PDF input and no layout. MarkItDown, anydoc and Markdownify MCP turn documents that already carry their text into Markdown: hand them a scan and MarkItDown sends it to an LLM, anydoc to its vendor's hosted service, while Markdownify MCP can't read it. Unstract and Receipt OCR go all the way to JSON: Unstract hands the reading to a text extractor and then an LLM, and Receipt OCR sends a photo of a receipt to a vision model.
OCR engine, vision model or converter: three families
A classic OCR engine returns the text of an image line by line, with coordinates. Tesseract.js ports the Tesseract engine to WebAssembly without changing it; PP-OCRv6, PaddleOCR's default OCR model, comes in three sizes from 1.5 to 34.5 million parameters. Positioned text doesn't tell you where the headings are or which cell a number belongs to, and Tesseract.js reads the image as a single block of text by default. PaddleOCR adds PP-StructureV3 for that job: it splits the page into regions and returns Markdown or JSON.
A vision-language model reads the whole page and writes out its structure directly. Chandra returns Markdown, HTML or JSON with the position of more than 15 block types, and turns a flowchart into Mermaid. Falcon OCR, the 270-million-parameter model that Falcon-Perception serves, transcribes text, formulas as LaTeX and tables as HTML. PaddleOCR-VL pairs a layout model with a 0.9-billion-parameter vision model, and its docs warn that calling that model on its own won't reproduce the published scores and may hallucinate. Falcon-Perception's README puts Falcon OCR 1.5 at about a third the size of PaddleOCR-VL and a twentieth of Chandra 2.
A document converter does no OCR: it reads the text that a Word, Excel or PowerPoint file or a born-digital PDF already holds, and writes it out as Markdown for an LLM. Microsoft's MarkItDown is written in Python; Firecrawl's anydoc in Rust, with no model and no Python; Markdownify MCP serves MarkItDown to an AI agent over the Model Context Protocol. A scanned PDF has no text layer, so anydoc stops with a NeedsOcr error listing the pages concerned, and MarkItDown only reads it through its markitdown-ocr plugin, which sends each page to an LLM.
Unstract and Receipt OCR add one more step: they pull named fields out of a document as JSON, using an LLM. Unstract describes those fields with prompts, for any kind of document, and leaves OCR to whichever text extractor you plug in. Receipt OCR only handles shop receipts.
What each tool needs to run
Read from the repositories and documentation on 7 October 2026.
| Tool | Licence | What you run | Scanned pages | Output | Latest release |
|---|---|---|---|---|---|
| PaddleOCR | Apache-2.0, weights included | Python plus PaddlePaddle, Transformers or ONNX Runtime; an inference server for PaddleOCR-VL in production | yes, locally | positioned text, Markdown, JSON, DOCX | 3.7.0, 11 Jun 2026 |
| Chandra | Apache-2.0; weights under a modified OpenRAIL-M | Python 3.10, a 16 to 80 GB NVIDIA GPU for vLLM, 10.6 GB of weights | yes, locally | Markdown, HTML, JSON | 0.2.0, 18 Mar 2026 |
| Falcon-Perception | Apache-2.0, weights included | Python 3.10 and PyTorch on CUDA 12.8, or MLX on Apple Silicon | yes, locally; PDFs through the Docker image | text, LaTeX, HTML tables | 1.0.0, 7 Apr 2026 |
| Tesseract.js | Apache-2.0 | a browser or Node.js 16; core and language files from jsDelivr by default | images only, no PDF | text, positioned words, hOCR | 7.0.0, 15 Dec 2025 |
| MarkItDown | MIT; OCR plugin tied to PyMuPDF (AGPL-3.0) | Python 3.10 to 3.14 | through an LLM, with the markitdown-ocr plugin | Markdown | 0.1.8, 21 Sep 2026 |
| anydoc | MIT | Rust crate, Node.js 20 module, Python wheel or WebAssembly | no; an option sends the document to Firecrawl Parse | Markdown | 0.2.4, 27 Aug 2026 |
| Markdownify MCP | MIT | Node.js, Bun to install, Python and markitdown | no | Markdown | 1.1.0, 1 May 2026 |
| Unstract | AGPL-3.0 | 25 Docker services, 8 GB of free RAM; an LLM, embeddings, a vector database, a text extractor | through the configured extractor | JSON fields | 0.192.7, 7 Oct 2026 |
| Receipt OCR | MIT | Python 3.10, an API key for an OpenAI-compatible vision model | through the LLM, receipt photos only | JSON | 0.4.0, 6 Oct 2026 |
Hardware is where the three vision-model projects part ways. PaddleOCR-VL runs locally for validation, but its docs move production onto an inference server: vLLM, SGLang, FastDeploy, MLX-VLM or llama.cpp. vLLM needs compute capability 8.0 and CUDA 12.6 there, and the docs advise against T4 and V100 cards; on x64 CPUs, PaddlePaddle, Transformers and llama.cpp are supported. Chandra's launcher starts a vLLM container with sudo on the NVIDIA runtime, with presets from a 16 GB T4 to an 80 GB H100. Falcon-Perception pulls PyTorch for CUDA 12.8, which needs NVIDIA driver 570 or later, or MLX on an Apple Silicon Mac, where only the batch engine is available. Its code sizes the cache in memory tiers starting at 15 GiB, and one user with two RTX 3090s saw the server ask for 56 GiB at startup.
The other six tools need no GPU of their own. Tesseract.js loads a WebAssembly core of about 3.9 MB plus only the languages you ask for, 2.9 MB compressed for English. On a long-running Node.js server, its docs advise recreating workers, every 500 jobs for instance, since their memory only grows. MarkItDown and anydoc are libraries, with no service and no database. Unstract sits at the far end: its Compose file starts 25 services, including Postgres, Redis, RabbitMQ, MinIO and Qdrant, and its runner mounts the host's Docker socket.
What leaves your machine
The vision models and Tesseract.js read pages locally once the weights or language files are downloaded. Tesseract.js fetches those files from the jsDelivr CDN by default, along with the worker script and the WebAssembly core, so an offline app has to host them and set workerPath, corePath and langPath. Falcon-Perception pulls its models from Hugging Face on first run, and PaddleOCR's docs install PaddlePaddle from paddlepaddle.org.cn.
With the converters, sending data out takes an option or an extra. MarkItDown passes a whole image to an OpenAI-compatible client with an instruction to write a detailed caption, so what comes back is a description, not a transcription. Its audio transcription extra, which Markdownify MCP installs, sends the sound to Google's speech recognition API with a generic key that the SpeechRecognition library reserves for personal or testing use. anydoc's --ocr hosted option sends the whole document, not just the scanned pages, to Firecrawl Parse; the Rust crate never makes a network call.
Receipt OCR has no documented local mode: every receipt goes to the configured provider, and the docs only cover OpenAI, Gemini and Groq. Unstract can keep everything on the server, provided you start the community Unstructured container, which its launch script leaves out, and serve the LLM and embeddings locally, through Ollama for example. Its other extractors, LLMWhisperer and LlamaParse, are hosted APIs.
What each tool can read
Scanned PDFs go straight into PaddleOCR and Chandra, and into Unstract through its extractor. Falcon-Perception takes images; a multi-page PDF only goes through its Docker image. Tesseract.js accepts no PDF at all, and its FAQ suggests rendering each page to an image with PDF.js or muPDF first. Receipt OCR rejects any file that isn't an image.
Born-digital PDFs also go through the converters, with limits. MarkItDown extracts the text layer with pdfminer: a prose PDF comes out as plain text with no headings, and only pages carrying a table or a form come out as Markdown tables, via pdfplumber. anydoc reads text-layer PDFs with pdf-inspector, but an open issue (#173) reports multi-column tables squashed into a single column.
On a scanned page, tables and formulas only come out of the models. PP-StructureV3 returns tables with coordinates for every cell; Chandra tells tables, forms, equations and checkboxes apart; Falcon OCR writes tables as HTML and formulas as LaTeX. Tesseract.js stops at positioned text, with no Markdown and no table cells.
Handwriting shows up in Chandra's README, which demonstrates cursive and handwritten maths, and in Falcon-Perception's, which recommends its end-to-end OCR for handwritten notes. Tesseract.js's FAQ rules it out. On languages, PP-OCRv6 reads 50 with a single model and PaddleOCR-VL more than a hundred; Chandra claims over 90 and Tesseract.js over 100. Falcon OCR's Hugging Face card doesn't list its own.
Older Office formats split the converters. anydoc reads binary .doc, .ppt and .xls files, OpenDocument and RTF, none of which MarkItDown accepts; Unstract also takes DOC, PPT, ODT and ODP. MarkItDown, for its part, reads HTML, audio, Outlook messages and Jupyter notebooks, which aren't on anydoc's list.
Where the projects stand
PaddleOCR shipped six releases between January and June 2026 and none since: 3.7.0 came out on 11 June. Chandra is still on 0.2.0 from March 2026, released alongside Chandra 2, and one developer accounts for 75 of the 82 attributed commits. Falcon-Perception has a single release, 1.0.0 from April 2026, marked beta, and Falcon OCR 1.5 support only lives on its main branch. Its REST server has no documented authentication, and a public report of a flaw exploitable without authentication (#37, 30 September 2026) was still unanswered on 7 October. Tesseract.js ships roughly one major version a year, each with breaking API changes, and has had a single commit since v7 in December 2025; its npm package was downloaded 3.9 million times in the week of 28 September 2026.
MarkItDown is still on 0.1.x, labelled beta on PyPI, and released 0.1.8 on 21 September 2026. anydoc, created on 3 August 2026, shipped all fourteen of its releases within three weeks in August; it hasn't had a commit since 28 August, and 64 pull requests are waiting. Markdownify MCP's main branch hasn't moved since v1.1.0 on 1 May 2026, and reported vulnerabilities remain open, including an SSRF through http://[::1] (CVE-2025-65512). Unstract runs the other way: 128 releases since January 2026, still on 0.x, and a launch script that pulls latest by default. Receipt OCR, maintained by one person, put out a maintenance release, v0.4.0, on 6 October 2026, nearly a year after the previous one.
Weight licences and paid plans
The code of all nine projects is open source: Apache-2.0 for PaddleOCR, Chandra, Falcon-Perception and Tesseract.js, MIT for MarkItDown, anydoc, Markdownify MCP and Receipt OCR, and AGPL-3.0 for Unstract. A vision model's weights carry their own licence, though. PaddleOCR and Falcon-Perception publish theirs under Apache-2.0. Chandra's fall under an OpenRAIL-M licence modified by Datalab: free for personal use and research, and off limits for anything else once the user, their employer or their organisation has passed 2 million dollars in revenue the previous year or raised more than 2 million. They are also off limits to any organisation that competes with Datalab, whatever its size, and the restrictions extend to the model's output. MarkItDown is under MIT, but its markitdown-ocr plugin depends on PyMuPDF, released under AGPL-3.0 or a commercial licence from Artifex.
Datalab sells a hosted API that, according to the README, runs a version of Chandra more accurate than the open weights; its Team plan costs 400 dollars a month. Firecrawl charges one credit per page for Parse, anydoc's hosted OCR: 1,000 free credits a month, then from 16 dollars (billed annually) to 599 dollars a month. Unstract's cloud starts at 499 dollars a month for 5,000 pages, and its open-source edition has no SSO, no human review screen and no second-LLM check of the answers. MarkItDown plugs into two Azure services billed per call, Document Intelligence and Content Understanding. PaddleOCR's docs, for their part, announce a free API covering up to 20,000 pages of document parsing a day, from the PaddleOCR website, which redirects to Baidu AI Studio. Tesseract.js, Falcon-Perception, Markdownify MCP and Receipt OCR have no paid plan, though Receipt OCR needs an API key with credits at the LLM provider.
