Mistral AI Launches Pixtral Large 123B: Open-Weights Frontier Multimodal Model for Enterprise Document AI
An open-weights vision-language powerhouse engineered for multi-page PDF reasoning, high-resolution chart analysis, and dense diagram extraction.
Ethan Walker
Table of Contents
Mistral AI has officially released Pixtral Large 123B, a frontier-class open-weights multimodal vision-language model. Built on Mistral Large 2 with a purpose-built 1-billion parameter vision encoder, Pixtral Large is designed for enterprise document AI workloads that require parsing dense financial reports, multi-page scientific papers, high-resolution engineering schematics, and complex data tables across a full 128k token context window.
Architecture: Mistral Large 2 + Custom Vision Encoder
Pixtral Large combines Mistral's strongest language model backbone — Mistral Large 2 with 123B parameters — with a custom vision encoder trained specifically on heterogeneous document types. Unlike general-purpose multimodal models optimized for natural image understanding, Pixtral Large's encoder is fine-tuned on structured document layouts: table grids, form fields, watermarked PDFs, handwritten annotations, and technical vector diagrams. This specialization results in significantly better extraction accuracy on enterprise document types compared to general vision-language models.
Key Technical Capabilities
- 123B parameter model with 1B parameter dedicated vision encoder
- 128k context window: ingests full multi-page PDFs in a single API call
- Multi-image reasoning: compares and cross-references up to 20 images simultaneously
- Table and structured data extraction with row/column relationship preservation
- Handwritten text recognition integrated into the standard generation pipeline
- Open weights: downloadable for self-hosted deployment under the Mistral Research License
Benchmark Performance
Pixtral Large achieves state-of-the-art results on DocVQA (document visual question answering), scoring 94.2% — outperforming GPT-4V and Claude 3.5 Sonnet on structured document understanding benchmarks. On MMMU (Massive Multi-discipline Multimodal Understanding), it scores 69.4%, placing it competitively against models twice its size. Mistral emphasizes that these benchmarks were achieved as an open-weights model, making Pixtral Large the strongest openly available multimodal model for document intelligence at launch.
Enterprise Deployment Use Cases
The primary target workloads for Pixtral Large are financial document processing (earnings reports, SEC filings, balance sheets), legal document review (contracts, compliance documents, regulatory filings), technical documentation extraction (engineering specs, patent applications), and healthcare imaging support (radiology reports, pathology results with reference image context). Organizations running on-premises document AI pipelines can now deploy a frontier-class multimodal model without API data-sharing concerns.
Access & Availability
Pixtral Large 123B is available for download from Hugging Face (mistralai/Pixtral-Large-Instruct-2411) and accessible via Mistral's La Plateforme API at $2 per million input tokens and $6 per million output tokens. Azure AI and Google Cloud Vertex AI integrations are planned for Q1 2026.
Found this useful? Share it:
Prefer NeedAITool on Google SearchAI Overviews
See our verified benchmarks & AI tool comparisons more frequently on Google.

Ethan Walker
I’m a technology writer passionate about AI tools, automation, productivity software, and emerging SaaS platforms. I spend my time testing digital tools and breaking down complex technologies into practical insights that help businesses, creators, and professionals work smarter.
AI Tools Mentioned in This Post
A European AI leader providing highly efficient open-weight models like Mistral 7B, Mixtral 8x7B, and their flagship Mistral Large.
RunPod is a leading globally distributed GPU cloud and serverless computing platform engineered specifically for artificial intelligence workloads. It provides developers, AI researchers, and enterprises with on-demand access to top-tier NVIDIA GPUs (including H100, A100, L40S, and RTX 4090) at up to 80% lower cost than traditional legacy hyperscalers.
