Addis PulseStudio

Tencent's WeVisDoc-4B: A Single-Model Document Parser That Scores 95 on OmniDocBench

A 4-billion-parameter vision-language model turns a page image into structured Markdown with LaTeX and HTML tables in one pass. The 2B sibling is within 0.3 points, which changes the hardware math for a small studio.

3 min read715 words

What happened

Tencent published WeVisDoc-4B to Hugging Face on 16 September 2026. The 4-billion-parameter model, fine-tuned from Qwen3-VL-4B-Instruct, turns a single page image into structured Markdown with embedded LaTeX formulas and HTML tables, and posts a 95.38 overall score on OmniDocBench v1.6, first among the end-to-end parsers in its comparison set.

Context

Document parsing has historically been a pipeline: OCR, layout detection, formula recognition, table structure recovery, each a separate model bolted together. WeVisDoc collapses that chain into one vision-language model that sees the page and emits the full structured output in a single pass. The release ships in two sizes: 4B (fine-tuned from Qwen3-VL-4B-Instruct) and 2B (fine-tuned from Qwen3-VL-2B-Instruct). A technical report is at arXiv 2609.20423; code is at github.com/Tencent/WeVisDoc.

How it works

You feed it a page image; it returns Markdown. Inside that Markdown, mathematical formulas appear as LaTeX and tables as HTML, so downstream text tools receive clean, parseable structure rather than a flat string of OCR tokens. The model is a standard Hugging Face safetensors checkpoint, served through vLLM, and requires Python 3.10 or later. It handles English and Chinese page content.

On OmniDocBench v1.6 the sub-scores are: TextEdit distance 0.036, FormulaCDM 96.81, TableTEDS 92.95, TableTEDS_S 95.34, ROEdit 0.125. On PureDocBench the three tracks split the signal: Clean 79.81, Digital Degraded 77.74, Real Degraded 69.08. All figures are means over three inference runs. The 2B variant scores 95.06 and 73.86 on the same two benchmarks.

Our read

The headline — 95.38, first in its comparison set — is the least interesting part. What matters to a small studio is the 2B-versus-4B gap: 95.06 versus 95.38 on OmniDocBench. You lose 0.32 points and get half the parameters. The 2B model fits on hardware that cannot run the 4B, and for a one- or two-GPU studio that is the model you deploy.

The second thing the table hides: PureDocBench's "Real Degraded" track drops from 79.81 (Clean) to 69.08, a 10.73-point cliff. That is the distance between a crisp PDF render and a phone photo of a printed script. If your workflow starts with photographed storyboards or scanned reference sheets, you are operating in the 69 range, not the 95 range. And no latency, throughput, or maximum-resolution figure is published in the source.

What the release does not address: at 212 downloads and no documented Transformers, Diffusers, or ComfyUI integration, the operational friction is real. You stand up a vLLM server, write a thin HTTP wrapper, and benchmark on your own hardware before calling this production-ready.

What this changes

On Monday, if you have a single GPU with at least 8 GB of VRAM (the 4B model in fp16 occupies roughly that much), you can spin up vLLM, point an HTTP client at it, and pipe photographed or scanned scripts, storyboards, and reference sheets through it to get clean Markdown before feeding text into a script-writer or prompt pipeline. The 2B variant is the sensible choice on 6 GB VRAM; the 0.32-point accuracy loss is negligible for document parsing. You will need a small wrapper script or a custom ComfyUI node to bridge the vLLM endpoint into your pipeline, because no off-the-shelf integration is documented. Benchmark pages-per-second on your own hardware before routing production material through it.

License

Apache-2.0, as stated on the Hugging Face model card. Commercial use is permitted with no revenue ceiling or attribution-only restriction. This applies to the 4B variant; the 2B variant's licence is not independently confirmed in this source, so check its model card before building on it.

Key takeaways

  • WeVisDoc-4B converts a single page image into structured Markdown with LaTeX formulas and HTML tables in one forward pass, scoring 95.38 on OmniDocBench v1.6.
  • The 2B variant (95.06) is within 0.32 points of the 4B model, making it the practical choice for studios on smaller GPUs.
  • Accuracy drops 10.73 points on degraded real-world pages (69.08 vs 79.81 on Clean), so input quality matters more than model size.
  • Serving requires vLLM and Python 3.10+; no ComfyUI, Transformers, or Diffusers pipeline is documented, so a custom wrapper is needed.
  • Apache-2.0 licence permits unrestricted commercial use for the 4B variant; the 2B licence should be verified separately.

Sources

  1. tencent/WeVisDoc-4B — model on Hugging Face — tier 2
llmdocument-processingcomputer-vision

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260918T131143Z