Introducing a hybrid approach to using Document AI and GenAI
DocVLM: ABBYY's Purpose-Built Document Vision-Language Model is Here
Slavena Hristova
July 27, 2026
Adding a new document type to your automation pipeline has traditionally meant one thing: training data. You need labeled examples, time to annotate, and a model that has seen enough variations before it can reliably extract a single field. For well-established document types with large sample sets, that process works. For everything else, it has been a persistent bottleneck.
DocVLM changes that equation. ABBYY's purpose-built Document Vision-Language Model, trained and fine-tuned by ABBYY AI Labs, is now being rolled out through a controlled release, with General Availability expected across ABBYY's Document AI platform by the end of 2026. DocVLM brings generative AI capabilities directly into the intelligent document processing (IDP) pipeline, enabling zero-shot extraction from previously unseen documents, without replacing the deterministic foundation that enterprise automation depends on.
This article walks through what DocVLM is, how it fits within ABBYY's broader PHOENIX portfolio of AI models, what it can do today, and what is coming next.
Jump to:
What is PHOENIX, and why does it matter for document processing?
What can DocVLM do today? Zero-shot extraction explained
What is coming next? Auto-labeling and human review enhancement
How does DocVLM relate to prompt-based extraction?
The hybrid approach: expanding automation without abandoning reliability
Key takeaways
- DocVLM is ABBYY's customized, ABBYY-hosted Document Vision-Language Model, trained and fine-tuned specifically for business document understanding.
- DocVLM is part of PHOENIX, ABBYY's integrated portfolio of AI models optimized for document processing.
- Zero-shot extraction is DocVLM's first available capability, enabling field extraction from unseen document types with no templates or training data required.
- Upcoming capabilities, such as auto-labeling and enhanced human review, on the roadmap utilize the model to accelerate and extend document processing beyond simple extraction, delivering reduced setup effort and streamlined validation.
- DocVLM is currently in a controlled release and is on track for General Availability by end of 2026.
- DocVLM reinforces ABBYY's hybrid approach in applying the right technology at the right time, combining deterministic and generative AI to maximize accuracy, minimize token usage, and expand what automation can achieve.
What is PHOENIX, and why does it matter for document processing?
DocVLM is part of PHOENIX — ABBYY's integrated portfolio of AI models, all optimized specifically for document processing. It is not a single model. It is a governed, multi-model architecture that routes the right technique to each document and each field, depending on what that task actually requires.
The principle is straightforward: apply the right technology for the job.
| Document or field scenario | PHOENIX approach |
| Fixed layouts, compliance-critical fields | Deterministic rules for precision and repeatability |
| Variable layouts, handwriting, complex scripts | ML models (CNN and RoBERTa-based) across 200 and more languages |
| Layout-aware understanding, unseen document types | Vision-language models (DocVLM) |
| Contextual reasoning, summarization, enrichment | Generative LLMs, either via BYOM or ABBYY's own hosted DocVLM and DSLMs |
This is the architecture behind ABBYY's hybrid AI approach. Large language models (LLMs) are powerful, but documents are not just text. They have tables, multi-column layouts, forms, stamps, and structural relationships that a general-purpose language model was not built to navigate reliably. PHOENIX exists to give every part of the pipeline the AI it actually needs.
PHOENIX currently includes production-ready ML and language models (CNN and RoBERTa-based) for classification and extraction. DocVLM extends this portfolio with generative, multimodal capability fine-tuned for business documents. On the horizon: domain-specific Small Language Models (DSLMs), compact generative models trained for particular industries or document types, offering higher performance at lower cost compared to large general-purpose systems.
What is DocVLM?
DocVLM is ABBYY's customized and ABBYY-hosted Document Vision-Language Model. It was trained and fine-tuned by ABBYY AI Labs on business document structures and layouts, making it fundamentally different from general-purpose vision-language models and LLMs.
Three properties define DocVLM's design:
- Purpose-built for documents. DocVLM understands the visual and structural characteristics of business documents, not just the text. It can reason across tables, forms, headers, and spatial relationships the way a document expert would.
- ABBYY-hosted for data privacy. Unlike BYOM integrations that send data to external providers, DocVLM runs within ABBYY's infrastructure. Organizations processing sensitive or regulated content maintain full data control.
- Cost-efficient by design. Because DocVLM is optimized for document tasks rather than general reasoning, it operates more efficiently than frontier models and general-purpose VLMs, reducing token consumption and infrastructure cost.
DocVLM is callable as a model within the ABBYY Document AI platform (Vantage and FlexiCapture), slotting into the IDP pipeline at multiple stages: optical character recognition (OCR) and ICR, document classification, data augmentation, and human-in-the-loop review.
What can DocVLM do today? Zero-shot extraction explained
The first production capability available through DocVLM is zero-shot extraction.
Zero-shot extraction means exactly what it says: the model pulls structured fields from document types it has never seen before, with no templates, no labeled training data, and no configuration beyond specifying what you want to extract. You describe the field. DocVLM finds it.
This matters enormously for organizations dealing with document variety at scale. Most enterprise automation programs cover a handful of high-volume, well-structured document types reasonably well. The long tail — contracts, correspondence, supplier forms, niche industry documents — remains manual precisely because training a model for each type is expensive and slow.
Zero-shot extraction directly addresses that gap. It allows teams to bring new document types into automated workflows faster, without waiting for a training dataset to accumulate.








