Production applications that rely on large language models require more than a working API call. They demand careful attention to prompt design, cost control, output reliability, and the ability to connect models to external systems. This article covers the core engineering patterns for LLM API integration: how prompt engineering shapes output quality, how the request/response lifecycle works at the token level, how to minimize spend without degrading quality, how to use function calling to give models access to tools, and how to enforce structured outputs for downstream consumption. It also covers semantic search with vector databases and Retrieval-Augmented Generation (RAG), which grounds LLM responses in real data.
Prompt Engineering for LLM APIs
Prompt engineering is the practice of crafting instructions that guide a generative AI model (LLM or VLM) to produce a desired output. Think of it as programming with natural language instead of code. The right prompt bridges the gap between a user's intent and the model's response.
Why does this matter? A well-engineered prompt improves accuracy, relevance, and overall quality. Without it, you get generic, irrelevant, or factually incorrect answers. Specific instructions also unlock capabilities like summarization and code generation that a simple query cannot achieve. Precise prompts reduce iteration, enable personalized content (such as a marketing email in a specific tone), and can steer models away from biased or inappropriate outputs.
Building a Prompt Step by Step
A strong prompt follows a structured approach with up to seven components:
1. Choose the Role. Assign expertise, perspective, or style. The model adapts tone and depth to match. For example: "You are a classification expert capable of understanding sentiment from customer reviews."
2. Define the Goal. Be clear about the outcome. Without a specific goal, the model guesses, leading to vague or irrelevant answers. For example: "Your goal is to classify each customer review into positive or negative based on the tone and words they use."
3. Give Context. Provide instructions, background information, constraints, or examples. Context prevents hallucination and keeps responses relevant. Include few-shot examples when possible (e.g., "I loved the cheesecake" maps to positive, "The waiting time was too much for me!" maps to negative).
4. Specify Output Format. Define how the answer should be structured: table, JSON, list, or paragraph. This saves time and ensures consistency across responses.
5. Guide the Reasoning. Ask the model to explain its thought process. This prevents shallow answers, reveals the model's reasoning, and lets you verify correctness.
6. Control Scope (optional). Set boundaries on length, number of ideas, or level of detail. Without limits, responses can be overwhelming or too shallow. Constraints sharpen the output (e.g., "a short sentence, up to 15 words").
7. Iterate. Refine the prompt after seeing the first output. The first attempt is rarely perfect. Iteration helps you converge on the best answer through a refinement loop.
Here is a complete prompt that applies all of these steps:
PROMPT = """
You are a classification expert capable of understanding sentiment from customer's reviews.
Your goal is to classify each customer review into positive or negative based on the tone and words they use.
**Instructions:**
1. Read and analyze the customer review
2. Extract the sentiment ('positive' or 'negative') based on the words used and the tone of voice. For example:
- "I loved the cheesecake" -> positive
- "The waiting time was too much for me!" -> negative
3. Return a JSON with the following keys:
- 'sentiment' - the sentiment identified in the customer review ('positive' or 'negative')
- 'reasoning' - a short sentence (up to 15 words) with the reason for your classification into positive or negative.
"""
LLM API Architecture: Requests, Tokens, and Pricing
Every LLM API interaction follows the same fundamental pattern. The client sends an HTTP request containing a prompt (and optionally a system instruction, conversation history, or attached media). The provider's inference server processes the input tokens, generates output tokens autoregressively, and returns the completed response.
The Token Economy
Tokens are the atomic billing unit. A token roughly corresponds to 3-4 characters in English text. Both input and output tokens are metered, but at different rates. For most providers, output tokens cost significantly more than input tokens because each output token requires a separate forward pass through the model, whereas input tokens are processed in a single parallel pass.
A basic OpenAI API call illustrates this lifecycle:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5",
instructions="You are a sentiment classifier. Return JSON with keys: sentiment, reasoning.",
input="The food was delicious!!!"
)
print(response.output_text)
The equivalent call using Google's Gemini SDK:
from google import genai
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-pro",
contents="Classify the sentiment of: 'The food was delicious!!!'"
)
print(response.text)
Both APIs follow the same pattern: instantiate a client, specify a model, send input, receive generated text. The differences are in parameter naming and SDK conventions, not in the underlying architecture. These models represent the latest generation of LLMs, with significantly improved reasoning and multimodal capabilities compared to their predecessors.
Pricing Models
LLM providers publish per-token pricing that varies by model tier. OpenAI's pricing for GPT-5 family models shows a clear tiered structure:
| Model | Input / 1K tokens | Output / 1K tokens |
|---|---|---|
| gpt-5 | $0.00125 | $0.01 |
| gpt-5-mini | $0.00025 | $0.002 |
| gpt-5-nano | $0.00005 | $0.0004 |
The 25x price difference between gpt-5 and gpt-5-nano means model selection is the single largest lever for cost optimization. You can compare costs across providers using online calculators such as Helicone's LLM cost tool. Understanding this pricing structure is a prerequisite for any production deployment.
Cost Optimization Strategies for LLM API Integration
LLM cost optimization operates on four axes: model selection, token management, caching, and batching. Each compounds with the others.
Model Selection: Right-Size for the Task
Large models like GPT-5 and Gemini Pro are expensive and often unnecessary for straightforward tasks. Keyword extraction, text classification, summarization, and simple routing can be handled by smaller models like GPT-5-mini or Gemini-2.5-flash at a fraction of the cost.
The most effective pattern is a multi-tiered approach: use a small, fast model for initial filtering or classification, then escalate to a larger model only for tasks that require complex reasoning or generation. For example, a customer support pipeline might use gpt-5-nano to classify ticket urgency and route to the correct department, then invoke gpt-5 only for tickets that require a nuanced, multi-paragraph response.
from openai import OpenAI
client = OpenAI()
def classify_then_respond(ticket_text: str) -> str:
# Step 1: Cheap classification with nano model
classification = client.responses.create(
model="gpt-5-nano",
instructions="Classify this support ticket as 'simple' or 'complex'. Return only the label.",
input=ticket_text
)
label = classification.output_text.strip().lower()
# Step 2: Route to appropriate model
if label == "simple":
response = client.responses.create(
model="gpt-5-nano",
instructions="You are a helpful support agent. Provide a brief, accurate response.",
input=ticket_text
)
else:
response = client.responses.create(
model="gpt-5",
instructions="You are a senior support agent. Provide a thorough, empathetic response.",
input=ticket_text
)
return response.output_text
Token Management
Four techniques reduce token consumption directly:
Minimize input token usage. Write concise, structured prompts. Remove redundant instructions. Every unnecessary word in your system prompt is billed on every request.
Reuse context through summarization. For multi-turn conversations, periodically summarize the conversation history and replace the full message array with the summary. This prevents linear cost growth as conversations lengthen.
Set max_tokens limits. Restricting output length avoids runaway generation. If you need a one-sentence answer, set max_tokens to 50 rather than allowing the model to generate 500 tokens.
Use lower temperature for deterministic tasks. A temperature of 0 or near-0 produces shorter, more focused outputs for classification and extraction tasks. Higher temperature increases variability and often length.
response = client.responses.create(
model="gpt-5-mini",
instructions="Extract the product name from this review. Return only the product name.",
input=review_text,
max_output_tokens=30,
temperature=0.0
)
Caching and Batching
For applications that process similar or repeated queries, implement a semantic cache. Hash the input prompt and store the response. On subsequent identical (or near-identical) requests, return the cached response without making an API call.
Batching is relevant when processing large datasets. Rather than sending requests sequentially, use asynchronous processing with concurrency limits to maximize throughput while respecting rate limits:
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI()
semaphore = asyncio.Semaphore(10) # max 10 concurrent requests
async def process_item(item: str) -> str:
async with semaphore:
response = await client.responses.create(
model="gpt-5-nano",
instructions="Classify the sentiment as positive or negative.",
input=item
)
return response.output_text
async def process_batch(items: list[str]) -> list[str]:
tasks = [process_item(item) for item in items]
return await asyncio.gather(*tasks)
Function Calling: Connecting LLMs to External Tools
Function calling (also referred to as tool use) enables an LLM to invoke external functions by generating structured arguments. The model does not execute the function itself. Instead, it outputs a JSON object specifying which function to call and with what parameters. Your application code executes the function and optionally feeds the result back to the model for further reasoning.
This mechanism transforms an LLM from a text generator into a reasoning agent that can interact with databases, APIs, file systems, and any other programmatic interface. For a broader look at how tool integration standards are evolving, see our overview of the Model Context Protocol (MCP).
OpenAI Function Calling with the Responses API
OpenAI's function calling requires defining tools as JSON schemas. The model decides when to invoke a tool based on the user query and the tool descriptions:
from openai import OpenAI
import json
client = OpenAI()
# Define tools the model can use
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a given city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "The city name, e.g. 'London'"
},
"units": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature units"
}
},
"required": ["city"]
}
}
}
]
response = client.responses.create(
model="gpt-5",
input="What's the weather like in Lisbon?",
tools=tools
)
# The model returns a function call, not a text response
for output in response.output:
if output.type == "function_call":
args = json.loads(output.arguments)
print(f"Function: {output.name}, Args: {args}")
# Execute the actual function here
# weather_data = get_weather(**args)
Building an Agent with LangChain Tools
LangChain is a versatile framework for building applications with LLMs like GPT-5, Claude, Gemini, or Mistral. It goes beyond single-prompt interactions by offering tools (connecting LLMs to databases, APIs, files, and web pages), chains (structuring LLM output to feed into subsequent steps automatically), multi-cloud integration (GCP, AWS, Azure), and built-in abstractions for vector databases, document loading, chunking, and embedding generation.
LangChain provides a higher-level abstraction for function calling through its agent framework. You define Python functions as tools, and the agent automatically decides when to invoke them:
from langchain.tools import Tool
from langchain.agents import initialize_agent, AgentType
from langchain_google_genai import ChatGoogleGenerativeAI
# Define a retrieval function
def retrieve_documents(query: str) -> str:
"""Retrieve relevant documents to answer user questions.
Args:
query (str): user question
Returns:
str: relevant documents
"""
results = vectorstore.similarity_search_with_relevance_scores(query, k=2)
documentation = [result[0].page_content for result in results]
return documentation
# Wrap it as a LangChain tool
doc_tool = Tool(
name="RetrieveDocuments",
func=retrieve_documents,
description="Retrieves relevant documentation for answering the user's question."
)
# Initialize the LLM and agent
llm = ChatGoogleGenerativeAI(model="gemini-2.5-flash")
# Note: initialize_agent and AgentType.ZERO_SHOT_REACT_DESCRIPTION are legacy APIs.
# Current LangChain uses create_react_agent from langgraph.prebuilt.
agent = initialize_agent(
tools=[doc_tool],
llm=llm,
agent=AgentType.ZERO_SHOT_REACT_DESCRIPTION,
verbose=True,
handle_parsing_errors=True
)
result = agent.run("How does RAG work?")
The agent reads the tool's name, description, and the function's docstring to determine when and how to call it. It generates the correct arguments, executes the function, reads the output, and then formulates a final answer. This is function calling in action: the LLM produces structured output that maps to a method call.
Structured Outputs: Reliable JSON from LLMs
Free-form text generation is useful for conversational interfaces but problematic for automated pipelines. When an LLM's output needs to feed into a database, an API, or another processing step, you need guaranteed structure. Structured output enforcement ensures the model returns valid JSON conforming to a predefined schema.
Why Structured Outputs Matter
The core benefits are:
- Reliability and predictability. Free-form text varies in structure, format, and content across invocations. Schema enforcement makes responses machine-readable and predictable.
- Automation. Structured outputs can be directly consumed by downstream software without parsing heuristics.
- Improved error handling. With a defined schema, you can validate responses using type checking before passing them to the next pipeline stage.
- Tool integration. Agent frameworks (LangChain, AutoGPT, and other agentic frameworks) rely on structured outputs to trigger actions in external tools (APIs, databases, scripts). Free-form text is ambiguous; structured outputs map directly to method calls.
Using Pydantic with LangChain for Structured Output
The most robust approach combines Pydantic models (for schema definition) with LangChain's with_structured_output method:
from pydantic import BaseModel, Field
from typing import List
from langchain_google_genai import ChatGoogleGenerativeAI
class SentimentResult(BaseModel):
sentiment: str = Field(description="The sentiment: 'positive' or 'negative'")
confidence: float = Field(description="Confidence score between 0 and 1")
reasoning: List[str] = Field(description="List of reasons supporting the classification")
llm = ChatGoogleGenerativeAI(model="gemini-2.5-flash")
structured_llm = llm.with_structured_output(SentimentResult)
result = structured_llm.invoke("The product quality exceeded my expectations, but shipping was slow.")
print(result.sentiment) # "positive"
print(result.confidence) # 0.72
print(result.reasoning) # ["Product quality praised", "Shipping complaint is secondary"]
The Pydantic model serves double duty. The Field(description=...) annotations tell the LLM what each field should contain. The type annotations (str, float, List[str]) define the expected data types, and Pydantic validates the output at runtime. If the model returns malformed data, Pydantic raises a validation error rather than silently passing bad data downstream.
OpenAI Structured Output with JSON Schema
OpenAI also supports structured outputs natively by passing a JSON schema in the API call:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5",
instructions="Analyze the review and return structured data.",
input="Great battery life but the screen is too dim.",
text={
"format": {
"type": "json_schema",
"name": "review_analysis",
"schema": {
"type": "object",
"properties": {
"overall_sentiment": {
"type": "string",
"enum": ["positive", "negative", "mixed"]
},
"aspects": {
"type": "array",
"items": {
"type": "object",
"properties": {
"feature": {"type": "string"},
"sentiment": {"type": "string"}
}
}
}
},
"required": ["overall_sentiment", "aspects"]
}
}
}
)
This guarantees the response conforms to the specified schema at the API level, without requiring client-side parsing or retry logic.
Semantic Search and Vector Databases
Semantic search retrieves information by understanding the meaning of a query rather than just matching keywords. This is a foundational building block for RAG systems and tool-augmented LLM applications.
How Semantic Search Works
The process has three steps:
- Convert documents to embeddings. Transform text (documents, chunks) into numerical vector representations using an embedding model.
- Convert the user query to an embedding. The same embedding model transforms the query into the same vector space.
- Calculate similarity. Use a distance metric (cosine similarity, Euclidean distance) to retrieve the top-k documents most similar to the query.
Implementation with PGVector
A practical implementation involves chunking documents, generating embeddings, and storing them in a vector database like PGVector (a PostgreSQL extension):
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.vectorstores import PGVector
from langchain_openai import OpenAIEmbeddings
# 1. Example documents
documents = [
"LangChain makes it easier to build applications with LLMs.",
"Retrieval-Augmented Generation combines semantic search and LLMs.",
"Semantic search retrieves meaning-based matches instead of keyword matches."
]
# 2. Split documents into chunks (useful for long texts)
text_splitter = RecursiveCharacterTextSplitter(chunk_size=200, chunk_overlap=20)
docs = text_splitter.create_documents(documents)
# 3. Create embeddings using OpenAI
embeddings = OpenAIEmbeddings(model="text-embedding-ada-002")
# 4. Connection details for Postgres
CONNECTION_STRING = PGVector.connection_string_from_db_params(...)
# 5. Store documents in Postgres with pgvector
vectorstore = PGVector.from_documents(
documents=docs,
embedding=embeddings,
collection_name="my_collection",
connection_string=CONNECTION_STRING,
)
# 6. Semantic search query
query = "How does RAG work?"
results = vectorstore.similarity_search_with_relevance_scores(query, k=2)
Splitting documents into chunks is important for several reasons: it allows you to store entire PDFs in your vector database, retrieve only the parts relevant to a user query, reduce noise sent to the LLM (lowering cost and hallucinations), and maintain context between chunks through overlap. The choice of embedding model also matters significantly. Researching which model performs best for your specific use case can mean the difference between a good application and a mediocre one.
Retrieval-Augmented Generation (RAG)
RAG is a technique that combines a retriever (vector database) and a generator (LLM) to improve the response quality of the LLM. It uses the retriever during inference time to build a richer prompt by adding context and knowledge based on the most relevant documents for the user query.
RAG Architecture: Retriever and Generator
A RAG system has two main components:
The Retriever handles document ingestion and search. Documents are broken into chunks, converted to embeddings using an embedding model, and stored in a vector database. At query time, the retriever converts the user query into an embedding, performs semantic search, and returns the top-k most similar documents.
The Generator is the LLM responsible for producing the final answer. It receives both the user query and the retrieved documents, then generates a response grounded in that context.
Why Use RAG?
RAG provides three key advantages over standalone LLMs:
- Updatable knowledge. You can easily update what the system knows by replacing or adding documents to the vector database. No model retraining required.
- Explainability. Users can check which documents were retrieved to provide context, something not possible with a standalone LLM.
- Reduced hallucinations. By providing accurate, up-to-date information through the retrieved documents, RAG produces more factually grounded responses.
RAG Implementation
The implementation builds on the semantic search code above. The key addition is a prompt that instructs the LLM to answer based only on the retrieved documentation:
from openai import OpenAI
PROMPT = """
You are a helpful and friendly QA chatbot.
Your goal is to help the users by answering their questions.
**Instructions:**
1. Read and analyze the user question
2. Read and analyze the documentation provided
3. Answer the user question taking into account only the documentation provided
**Documentation:**
{documentation}
"""
client = OpenAI()
# Retrieve relevant documents via semantic search
results = vectorstore.similarity_search_with_relevance_scores(query, k=2)
documentation = '\n'.join([result[0].page_content for result in results])
# Send documents and query to LLM
resp = client.responses.create(
model="gpt-5",
instructions=PROMPT.format(**{"documentation": documentation}),
input=query
)
Multimodal API Capabilities
Modern LLM APIs accept more than text. GPT-5 and Gemini process images alongside text in a unified inference pass, enabling applications that reason across modalities. Instead of relying on separate tools for image generation and text analysis, these models integrate both capabilities. This unified intelligence allows models to reason across modalities rather than processing them in isolation.
Vision: Image Understanding
To send an image to the API, encode it as base64 (because most REST APIs accept only text-based data formats like JSON, which cannot directly include binary image data) and include it in the request:
import base64
from openai import OpenAI
def load_image(image_path: str) -> str:
with open(image_path, "rb") as f:
return base64.b64encode(f.read()).decode("utf-8")
client = OpenAI()
base64_image = load_image("product_photo.png")
response = client.chat.completions.create(
model="gpt-5",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Generate a product title and description from this image."},
{
"type": "image_url",
"image_url": {
"url": f"data:image/png;base64,{base64_image}"
}
}
],
}
],
)
The Gemini equivalent uses Part.from_bytes:
from google import genai
base64_image = load_image("product_photo.png")
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[
genai.types.Part.from_bytes(
data=base64_image,
mime_type="image/png",
),
"Generate a product title and description based on this image."
]
)
Practical multimodal applications span multiple industries:
- Retail and e-commerce. Product catalog automation: upload product images and the model generates titles, descriptions, and tags.
- Finance. Document processing: extract insights from scanned reports, charts, or presentations.
- Legal and compliance. Cross-reference legal documents with photographic evidence.
- Education. Explain a graph by talking through it.
- Medicine. Analyze an X-ray alongside patient notes.
- Image classification and OCR. Classify images into categories or extract structured data from invoices and receipts (name, company, value, VAT, address).
Image Generation and Editing
For image generation, OpenAI provides DALL-E 3 and Gemini offers native generation:
# OpenAI DALL-E 3
client = OpenAI()
response = client.images.generate(
model="dall-e-3",
prompt="A futuristic city skyline at sunset, photorealistic style",
size="1024x1024",
n=1
)
# Gemini image generation
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-flash-image-preview",
contents=["A futuristic city skyline at sunset, photorealistic style"],
)
Gemini also supports image editing, where you send an existing image along with a text instruction to modify it (e.g., "Create a pencil sketch image of this sneaker"). This uses the same generate_content API with both the image bytes and the edit prompt, plus a GenerateContentConfig that specifies both TEXT and IMAGE response modalities.
Production Scaling: Rate Limits, Error Handling, and Parallel Processing
Moving from prototype to production requires handling the realities of distributed systems: rate limits, transient failures, and throughput requirements.
Rate Limiting and Retry Logic
All LLM providers enforce rate limits (requests per minute, tokens per minute). A production client must handle HTTP 429 responses with exponential backoff:
import time
import random
from openai import OpenAI, RateLimitError
client = OpenAI()
def call_with_retry(prompt: str, max_retries: int = 5) -> str:
for attempt in range(max_retries):
try:
response = client.responses.create(
model="gpt-5-mini",
input=prompt
)
return response.output_text
except RateLimitError:
wait_time = (2 ** attempt) + random.uniform(0, 1)
time.sleep(wait_time)
raise Exception("Max retries exceeded")
Streaming for Low-Latency Applications
For user-facing applications, streaming reduces perceived latency by delivering tokens as they are generated rather than waiting for the complete response:
from openai import OpenAI
client = OpenAI()
stream = client.responses.create(
model="gpt-5-mini",
input="Explain how transformers work in three paragraphs.",
stream=True
)
for event in stream:
if hasattr(event, "delta"):
print(event.delta, end="", flush=True)
Parallel Processing with Concurrency Control
For batch workloads, concurrent processing with semaphore-based throttling maximizes throughput while staying within rate limits:
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI()
async def process_batch(
items: list[str],
model: str = "gpt-5-nano",
max_concurrent: int = 10
) -> list[str]:
semaphore = asyncio.Semaphore(max_concurrent)
results = []
async def process_single(item: str) -> str:
async with semaphore:
response = await client.responses.create(
model=model,
input=item
)
return response.output_text
tasks = [process_single(item) for item in items]
results = await asyncio.gather(*tasks, return_exceptions=True)
return [r if isinstance(r, str) else str(r) for r in results]
Monitoring and Observability
Track three metrics for every production LLM integration:
- Token usage per request. Monitor both input and output tokens to detect prompt bloat or unexpected generation patterns.
- Latency distribution. P50, P95, and P99 latency inform model selection and timeout configuration.
- Error rates by type. Distinguish between rate limit errors (capacity planning), validation errors (prompt engineering), and server errors (provider reliability).
Log the model name, token counts, latency, and cost for every API call. This data drives optimization decisions: if 80% of your spend comes from a single endpoint using gpt-5, and the task is classification, switching to gpt-5-nano could reduce that line item by 95%.
Key Takeaways
LLM API integration is an engineering discipline with concrete optimization levers. Prompt engineering, following a structured approach from role definition through iteration, directly determines output quality. Model selection alone can reduce costs by orders of magnitude. Function calling transforms LLMs from text generators into tool-using agents. Structured outputs eliminate the fragility of free-form text parsing. Semantic search and RAG ground LLM responses in real, retrievable data, reducing hallucinations and enabling updatable knowledge. Production deployment demands the same rigor as any distributed system: retry logic, concurrency control, and observability.
These patterns compose naturally. A production agent might use a nano model for intent classification, function calling to retrieve documents from a vector database, structured outputs to guarantee response format, and streaming to deliver results to the user with minimal perceived latency. Each technique addresses a specific failure mode or cost driver, and together they form the foundation of reliable LLM-powered applications.