Large language models spent their first years operating on a single modality: text. GPT-3, BERT, and their descendants processed sequences of tokens derived exclusively from written language. That constraint is now dissolving. The current generation of foundation models accepts images, audio, and video alongside text, reasons across those inputs jointly, and produces outputs that reference information from any combination of modalities. This shift builds directly on the rapid evolution of LLM architectures over the past several years.
This article covers the technical foundations behind multimodal AI, surveys the leading vision-language models, and identifies concrete applications relevant to data professionals working in production environments.
From Text-Only to Multimodal: Why Modalities Are Merging
The shift toward multimodal models is driven by a straightforward observation: most real-world information is not purely textual. Medical records contain scans. Financial reports embed charts. Manufacturing quality control relies on photographs. Customer support interactions include screenshots. A model that can only read text must depend on separate, disconnected systems to handle these other data types.
Early approaches to multimodal processing used pipeline architectures. An image classifier produced a label, which was then passed as text to a language model. This two-stage approach introduced information loss at every handoff. The image classifier might detect "a chart" but discard the axis labels, data values, and color encodings that carry the actual meaning.
Modern multimodal LLMs eliminate this bottleneck. They process raw pixels, audio waveforms, and text tokens through a unified architecture that maintains cross-modal relationships throughout the inference pass. The result is a model that can answer questions about the content of an image, transcribe and summarize audio, or describe what happens in a video within a single forward pass.
The Context Window Factor
One key enabler of multimodal convergence is the expansion of context windows. Processing an image requires significantly more tokens than processing a paragraph of text. A single high-resolution image can consume thousands of tokens. The progression from GPT-3.5-turbo's 16K token window to Gemini 1.5 Pro's 1M token window created the space needed to include visual and audio content alongside textual prompts.
Context window sizes across models illustrate this expansion:
| Model | Context Window (tokens) |
|---|---|
| GPT-3.5-turbo | 16,385 |
| Mistral-7B | 32,000 |
| Gemini 1.0 Pro | 32,000 |
| Claude 1 | 100,000 |
| GPT-4 Turbo | 128,000 |
| Claude 2.1 | 200,000 |
| Gemini 1.5 Pro | 1,000,000 |
Larger context windows carry trade-offs. Input token cost scales with context length, and benchmarks such as BABILong show that accuracy on factual retrieval tasks degrades as context grows. On single-fact QA, GPT-4o scored 100% at 0K context but dropped to 27% at 128K. GPT-4 Turbo showed a similar pattern, falling from 100% at 0K to just 6% at 128K. Beyond accuracy, long contexts also increase latency and make the system impractical for cost-sensitive applications. These degradation curves explain why retrieval-augmented generation (RAG) remains relevant even as context windows expand. RAG addresses several fundamental LLM limitations: the knowledge cutoff problem (models only know what they saw during training), the lack of source attribution, and the cost of stuffing entire document collections into a single prompt.
Vision-Language Models: Architecture and Capabilities
The leading multimodal LLMs, including GPT-4V, Gemini, and Claude 3, share a common architectural pattern. Each modality gets a dedicated encoder that projects raw inputs into a shared embedding space. A decoder (typically a transformer) then operates over the combined representation.
The Encoder-Decoder Pattern
A multimodal model processes input through modality-specific encoders: one for text tokens, one for image patches, and potentially one for audio spectrograms. Each encoder produces a sequence of vectors in a common high-dimensional space. The decoder attends over all of these vectors simultaneously, which allows it to generate text that references information from any input modality.
# Conceptual architecture of a multimodal model
# Each modality has its own encoder projecting to a shared space
class MultimodalModel:
def __init__(self):
self.text_encoder = TransformerEncoder(vocab_size=32000, dim=4096)
self.image_encoder = VisionTransformer(patch_size=16, dim=4096)
self.audio_encoder = AudioTransformer(dim=4096)
self.decoder = TransformerDecoder(dim=4096, layers=32)
def forward(self, text_tokens, image_patches, audio_frames):
text_embeds = self.text_encoder(text_tokens)
image_embeds = self.image_encoder(image_patches)
audio_embeds = self.audio_encoder(audio_frames)
# Concatenate all modality embeddings
combined = concat([text_embeds, image_embeds, audio_embeds])
# Decoder attends over all modalities jointly
output = self.decoder(combined)
return output
This design means the model does not treat an image as a separate entity from the text prompt. During attention computation, every token in the output sequence can attend to both the text instructions and the visual content. This is what enables tasks like "describe the trend shown in this chart" or "what text appears in the top-left corner of this screenshot."
GPT-4o Vision
OpenAI's GPT-4o accepts images alongside text prompts through the Chat Completions API. Images can be provided as URLs or base64-encoded data. The model handles photographs, diagrams, charts, screenshots, and handwritten text. It supports multiple images in a single request, enabling comparative analysis.
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "What data trend does this chart show?"},
{
"type": "image_url",
"image_url": {"url": "https://example.com/chart.png"},
},
],
}
],
max_tokens=500,
)
print(response.choices[0].message.content)
Gemini Multimodal
Google's Gemini family was designed as natively multimodal from the ground up. Gemini 1.5 Pro processes text, images, audio, and video within a single model, with context windows reaching 1M tokens. This capacity allows ingestion of entire video files. A 2-minute video can be processed by downsampling frames, detecting scenes, selecting key clips, and passing the result through a vision-language model.
Claude Vision
Anthropic's Claude 3 family (Haiku, Sonnet, Opus) accepts images alongside text. Claude processes images by encoding them into its token space, where they participate in the same attention mechanism as text tokens. The model handles document analysis, chart interpretation, and visual question answering.
Image Understanding: What Models Can and Cannot See
Multimodal LLMs perform well on a specific set of visual tasks and fail predictably on others. Understanding these boundaries is critical for building reliable applications.
Tasks where current models excel:
- Optical character recognition (OCR) on printed and handwritten text
- Chart and graph interpretation (bar charts, line plots, scatter plots)
- Object identification and scene description
- Screenshot and UI element recognition
- Document layout analysis (tables, headers, paragraphs)
Tasks where current models struggle:
- Precise spatial reasoning ("is object A to the left or right of object B" at small scales)
- Counting objects in dense scenes
- Reading low-resolution or heavily compressed images
- Understanding temporal sequences from a single frame
- Fine-grained visual distinctions (differentiating similar species, materials)
Two Strategies for Image Processing in RAG Systems
When integrating images into retrieval-augmented generation pipelines, two primary strategies exist:
1. Cross-modal embeddings. A multimodal embedding model maps both text and images into a single vector space. A query expressed in text can retrieve relevant images, and an image query can retrieve related text passages. The training objective aligns textual and visual representations so that a photo of a cat and the phrase "a domestic cat" produce similar vectors.
# Cross-modal embedding: text and images share a vector space
from sentence_transformers import SentenceTransformer
# Simplified example -- actual CLIP API may require different image handling
model = SentenceTransformer("clip-ViT-B-32")
# Encode text and image into the same space
text_embedding = model.encode("a photograph of a circuit board")
image_embedding = model.encode(Image.open("circuit_board.jpg"))
# Cosine similarity enables cross-modal retrieval
similarity = cosine_similarity(text_embedding, image_embedding)
2. Image-to-text summarization. A vision-language model generates a textual description of each image. That description is then embedded using a standard text embedding model and stored in the vector database alongside text chunks. At retrieval time, the system operates entirely in text space. This approach sacrifices some visual nuance but avoids the need for multimodal embedding infrastructure.
The choice between these strategies depends on the application. Cross-modal embeddings preserve visual information more faithfully. Image-to-text summarization is simpler to implement and integrates directly with existing text-based RAG pipelines.
Embedding Model Challenges
Multimodal embedding models face several practical challenges that affect production deployments:
- Context sensitivity. The same word or image region can carry different meanings depending on surrounding context. Embedding models must capture these nuances, which becomes harder as the input domain broadens.
- Scalability. Managing large-scale vector indices across multiple modalities requires careful infrastructure planning. Index sizes grow quickly when storing embeddings for images alongside text.
- Computational costs. Training and serving multimodal embedding models demands significantly more compute than text-only models, particularly for high-dimensional visual embeddings.
Chunking and Retrieval for Multimodal Content
When building RAG systems that handle mixed content, chunking strategy matters. Three common approaches apply:
- Fixed-size chunking. Split content into uniform segments with some overlap. Simple to implement but can break semantic units.
- Specialized chunking. Adapt chunk boundaries to the structure of the input (e.g., splitting on section headers, table boundaries, or image captions).
- Semantic chunking. Use embedding similarity to determine where to split, preserving meaning within each chunk.
After chunking, a two-stage search-and-rerank pipeline improves retrieval quality. The initial search retrieves hundreds of candidates by vector similarity, then a reranking model (such as those from Cohere or cross-encoder models on HuggingFace) narrows results to the most relevant handful of contexts.
Audio Processing: Speech, Sound, and Language
Audio enters multimodal systems through two main paths: transcription-first pipelines and native audio understanding.
Transcription-First Pipeline
The most established approach converts audio to text as a preprocessing step. Automatic speech recognition (ASR) systems like Whisper produce transcripts, which are then chunked, embedded, and stored in a vector database. At query time, the system retrieves relevant transcript segments and passes them to an LLM for answer generation.
import whisper
# Step 1: Transcribe audio
model = whisper.load_model("large-v3")
result = model.transcribe("meeting_recording.mp3")
transcript = result["text"]
# Step 2: Chunk the transcript
chunks = chunk_text(transcript, chunk_size=512, overlap=50)
# Step 3: Embed and store in vector database
embeddings = embedding_model.encode(chunks)
vector_db.upsert(embeddings, metadata=chunks)
# Step 4: Query and generate
query = "What decisions were made about the Q3 budget?"
relevant_chunks = vector_db.search(query, top_k=5)
answer = llm.generate(query, context=relevant_chunks)
This pipeline is mature and production-ready. Whisper achieves word error rates below 5% on English speech in controlled environments. The limitations are predictable. Accuracy degrades with background noise, overlapping speakers, heavy accents, and domain-specific terminology.
Native Audio Understanding
Gemini 1.5 Pro and GPT-4o process audio natively without an explicit transcription step. The audio waveform is tokenized directly and fed into the model alongside text tokens. This preserves information that transcription discards, including speaker tone, emphasis, background sounds, music, and non-speech audio events.
Native audio processing enables tasks that transcription pipelines cannot handle, such as identifying the emotional tone of a speaker, detecting environmental sounds, or answering questions about music content.
Cross-Modal Applications for Data Professionals
Combining modalities enables application categories that were previously impractical.
Document Analysis
Enterprise documents rarely contain only text. Annual reports, technical manuals, and research papers combine prose, tables, figures, and diagrams. A multimodal model can process an entire PDF page as an image, extracting information from both the text and the embedded visuals in a single pass.
import anthropic
import base64
client = anthropic.Anthropic()
# Read PDF page as image
with open("annual_report_page_15.png", "rb") as f:
image_data = base64.standard_b64encode(f.read()).decode("utf-8")
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": image_data,
},
},
{
"type": "text",
"text": "Extract all revenue figures from this page, "
"including those shown in the chart.",
},
],
}
],
)
Video Understanding
Video analysis decomposes into frame selection and per-frame reasoning. A typical pipeline for a 2-minute video operates as follows:
- Downsampling: reduce 3,600 frames (at 30 fps) to approximately 480 frames
- Scene detection: identify 10-15 distinct scenes
- Key clip detection: select approximately 100 representative frames
- Frame selection: choose 40 frames for VLM processing
- VLM analysis: extract textual descriptions from selected frames
The audio track is processed separately through ASR, producing a transcript that is time-aligned with the visual content. Both streams feed into a chunking and ingestion pipeline that stores the combined representation in a vector store for retrieval.
Visual Question Answering at Scale
Production visual QA systems combine multiple retrieval modalities. A query about a video might decompose into sub-queries. An ASR text query finds relevant dialogue, an object detection query locates specific items in frames, and an OCR query extracts on-screen text. The system merges results from these retrieval paths before generating a final answer.
Practical Integration: Building a Multimodal RAG Pipeline
A production multimodal RAG system extends the standard text-based RAG architecture with additional ingestion paths for images and audio.
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings
from langchain_core.documents import Document
def ingest_multimodal_document(pdf_path):
"""Process a document containing text, images, and tables."""
# Extract components
text_chunks = extract_text_chunks(pdf_path)
images = extract_images(pdf_path)
tables = extract_tables(pdf_path)
documents = []
# Process text chunks directly
for chunk in text_chunks:
documents.append(Document(
page_content=chunk,
metadata={"source": pdf_path, "type": "text"}
))
# Summarize images using a VLM, then store summaries
for img in images:
summary = vlm.describe_image(img)
documents.append(Document(
page_content=summary,
metadata={"source": pdf_path, "type": "image_summary"}
))
# Summarize tables, then store summaries
for table in tables:
summary = vlm.describe_table(table)
documents.append(Document(
page_content=summary,
metadata={"source": pdf_path, "type": "table_summary"}
))
# Embed and store all documents
vectorstore = Chroma.from_documents(
documents,
OpenAIEmbeddings(),
collection_name="multimodal_docs"
)
return vectorstore
This approach uses image-to-text summarization as the bridge between visual content and the text-based vector store. For applications requiring higher visual fidelity, replace the text embedding model with a multimodal embedding model (such as CLIP) and store raw image embeddings alongside text embeddings.
For domains where relationships between entities matter more than raw text similarity, GraphRAG offers an alternative. Instead of storing flat text chunks, GraphRAG uses an LLM to extract entities and relationships from documents, then stores them in a graph database. At query time, the system performs both vector similarity search and graph traversal to retrieve relevant context. This approach works well for knowledge-intensive tasks where connections between concepts (people, organizations, events) are central to answering questions accurately.
Current Limitations and the Path Forward
Multimodal AI has advanced significantly, but several constraints affect production deployments.
Hallucination across modalities. Models sometimes describe objects or text that do not appear in the input image. This is particularly dangerous in high-stakes applications like medical imaging or legal document review. Validation pipelines that cross-check model outputs against the source material remain essential.
Cost and latency. Processing images and audio consumes substantially more tokens than equivalent text. A single image may cost 1,000+ tokens. For applications processing thousands of documents per day, the token cost of multimodal processing can be 5-10x higher than text-only alternatives.
Spatial reasoning gaps. Current models perform poorly on tasks requiring precise geometric understanding, including measuring distances in images, understanding 3D spatial relationships, or performing exact counting in cluttered scenes.
Audio quality sensitivity. While native audio models handle clean speech well, performance degrades with background noise, cross-talk, and non-standard accents. Production systems still benefit from audio preprocessing and enhancement before model ingestion.
Evaluation complexity. Evaluating multimodal outputs is harder than evaluating text-only outputs. Statistical metrics like BLEU (which measures n-gram precision with a brevity penalty) and ROUGE (which measures recall of n-gram overlap) operate on text pairs and do not capture visual understanding. Model-based metrics like BERTScore (which uses BERT embeddings to measure semantic similarity between generated and reference text) and G-Eval (which uses an LLM with chain-of-thought to score outputs) offer more nuanced evaluation but add cost and latency. Evaluating whether a model correctly interpreted a chart still requires domain-specific validation logic that does not generalize across tasks.
What to Expect Next
The trajectory points toward models that handle more modalities with less preprocessing. Native video understanding (processing raw video rather than sampled frames) is an active research area. Real-time multimodal interaction (a model simultaneously processing a live camera feed, microphone input, and text chat) is moving from research demonstrations to product features.
For data professionals, the practical implication is clear: pipelines that currently operate on text-only data will increasingly need to handle images, audio, and video. Investing in multimodal ingestion infrastructure, cross-modal embedding strategies, and evaluation frameworks for non-text outputs positions teams to adopt these capabilities as they mature.
The models are no longer text-only. The data pipelines should not be either.