Prompt engineering is the process of designing and optimizing text inputs to deliver consistent, high-quality responses from large language models for a given application objective. The field has matured from simple instruction-writing into a structured discipline with a taxonomy of techniques, each suited to different task types and complexity levels. This reference covers the core techniques, ranks them by empirical effectiveness, and provides a decision framework for selecting the right approach.
The Prompting Taxonomy: Basic to Advanced
Prompting techniques organize into three tiers based on complexity and the type of reasoning they elicit:
Basic Techniques operate directly on the model's pretrained knowledge with minimal scaffolding:
- Zero-shot prompting
- Few-shot prompting
- Chain-of-thought (CoT) prompting
Intermediate Techniques introduce structured exploration or aggregation over multiple outputs:
- Self-consistency (majority voting over CoT paths)
- Tree of Thoughts (hierarchical branching through solution spaces)
Advanced Techniques combine reasoning with external data or tool access:
- RAG (Retrieval-Augmented Generation)
- ART (Automatic Reasoning and Tool-use)
- ReAct (Reasoning and Action combined)
Each tier builds on the previous one. A zero-shot prompt is the atomic unit; few-shot adds exemplars; CoT adds explicit reasoning steps; self-consistency and Tree of Thoughts add multiple parallel reasoning paths; ReAct and ART add external grounding through actions and tools.
| Technique | Description | Best Use Case |
|---|---|---|
| Zero-shot | Direct task, no examples | Simple generation, classification |
| Few-shot | In-context demo pairs | Small dataset annotation, code suggestion |
| Chain-of-thought | Step-by-step logic | Data wrangling, analytical workflows |
| ReAct / ToT | Mixed reasoning and retrieval | Fact-based QA, agentic data tasks |
| RAG | Retrieval-augmented generation | Large database/document analysis |
Zero-Shot Prompting: When Task Knowledge Is Sufficient
Zero-shot prompting provides the model with only an instruction and no examples. It relies entirely on the model's pretrained capabilities and instruction tuning. This works when the task is well-defined, the expected output format is standard, and the domain does not require specialized patterns the model has not seen during training.
Zero-shot is the correct starting point for any new task. If it achieves acceptable accuracy (roughly 80%+ for production use cases), additional complexity adds cost without proportional benefit.
Example: Zero-shot classification
Classify the following customer review as "positive", "negative", or "neutral".
Review: "The delivery was fast but the packaging was damaged and two items were missing."
Classification:
Example: Zero-shot extraction
Extract all dates mentioned in the following text. Return them in ISO 8601 format.
Text: "The contract was signed on March 15, 2024 and expires December 31, 2025."
Dates:
Zero-shot works best for tasks where the instruction itself is unambiguous and the output space is constrained. Classification, extraction, translation, and simple summarization are strong candidates. It fails when the output format is complex, when the task requires domain-specific conventions, or when edge cases dominate.
Few-Shot Prompting: Learning from In-Context Examples
Few-shot prompting provides 2-5 input-output pairs before the actual query. The model uses these examples to infer the task pattern, output format, and handling of edge cases. According to effectiveness rankings from the Prompt Report (Schulhoff et al.), few-shot prompting ranks as the highest-effectiveness technique overall, outperforming even chain-of-thought in broad benchmarks.
Six factors govern the quality of few-shot examples:
- Quantity: 2-5 examples typically yield optimal results. Beyond 5, returns diminish and context window costs increase.
- Order: Place the most representative example first. Models exhibit primacy bias.
- Label distribution: Balance across categories. Skewed examples produce skewed outputs.
- Quality: Each example must be correct and unambiguous. One bad example degrades the entire set.
- Format: Use consistent structure across all examples. XML tags or Q&A format both work. Do not mix formats.
- Similarity: Examples should be semantically close to the expected input distribution.
Example: Few-shot sentiment classification with edge cases
Classify the sentiment of each tweet as "Positive" or "Negative".
<example>
<input>Great game tonight!</input>
<output>Positive</output>
</example>
<example>
<input>Oh yeah, I really needed a flight delay tonight! Excellent!</input>
<output>Negative</output>
</example>
<example>
<input>The new update fixed all my issues, finally.</input>
<output>Positive</output>
</example>
Now classify:
<input>Sure, that was the best movie I've seen since Plan 9 from Outer Space.</input>
<output>
The second and fourth examples above are critical: they demonstrate sarcasm detection, teaching the model that surface-level positive language can carry negative sentiment. Including both positive and negative examples that share lexical features forces the model to reason about intent rather than keyword-match.
A common mistake is providing only "happy path" examples. Production inputs contain sarcasm, negation, ambiguity, and domain jargon. Few-shot examples should cover these failure modes explicitly.
Chain-of-Thought Prompting: Explicit Step-by-Step Reasoning
Chain-of-thought (CoT) prompting instructs the model to produce intermediate reasoning steps before arriving at an answer. Introduced by Wei et al. (2022), CoT activates the model's capacity for multi-step inference, which standard prompting bypasses.
CoT exists in two variants:
Zero-shot CoT appends a reasoning trigger to the prompt:
A store has 15 apples. It sells 8 in the morning and receives a shipment of 12 in the afternoon.
A customer then buys 6. How many apples does the store have at the end of the day?
Think step by step.
The model produces:
Step 1: Start with 15 apples.
Step 2: Sell 8 in the morning: 15 - 8 = 7.
Step 3: Receive shipment of 12: 7 + 12 = 19.
Step 4: Customer buys 6: 19 - 6 = 13.
The store has 13 apples.
Few-shot CoT provides examples that include the reasoning chain, not just the final answer:
Q: A restaurant serves 45 customers at lunch and 62 at dinner. Each customer generates
an average of $18 in revenue at lunch and $32 at dinner. What is total daily revenue?
A: Let me work through this step by step.
Lunch revenue: 45 customers x $18 = $810
Dinner revenue: 62 customers x $32 = $1,984
Total daily revenue: $810 + $1,984 = $2,794
Q: A warehouse processes 120 orders per hour with 3 workers. They need to process
2,400 orders. Each additional worker increases throughput by 40 orders/hour. How many
workers are needed to finish in 8 hours?
A:
Empirical data from the Prompt Report shows that few-shot CoT consistently delivers superior results compared to all other prompting variants tested, with accuracy scores around 0.692 versus 0.627 for plain zero-shot. The combination of examples plus explicit reasoning chains gives the model both the pattern to follow and the reasoning depth to handle novel inputs.
CoT is most effective for arithmetic, logic, multi-step reasoning, data analysis workflows, and any task where intermediate computations affect the final answer. It adds latency and token cost due to the longer output, so apply it selectively.
Decomposition: Breaking Problems into Subproblems
Decomposition prompting explicitly breaks down a complex problem into smaller, more manageable subproblems before attempting solutions. Unlike chain-of-thought, which reasons step by step within a single problem context, decomposition restructures the problem itself.
The key prompt pattern is: "Before answering, tell me some subproblems that would need to be solved first." This forces the model to identify dependencies and tackle each piece independently, reducing errors on tasks with multiple interacting constraints.
Example: Decomposition prompt
I need to build a customer churn prediction pipeline.
Before building it, list the subproblems that need to be solved first,
then address each one.
Decomposition ranks as the second-highest effectiveness category in the Prompt Report, just behind few-shot prompting. It is particularly valuable for troubleshooting, decision-making, critical thinking, and any situation where multiple angles need consideration.
Self-Criticism: Iterative Refinement Through Self-Review
Self-criticism prompting asks the model to critique and refine its own responses. The pattern follows a four-step process: initial task completion, self-review, feedback analysis, and revised response.
A practical prompt structure: "Check your response and confirm it's correct" or "Great criticism, why don't you go ahead and implement that?" This encourages the model to identify gaps, errors, or improvements in its own output before presenting the final answer.
Self-criticism ranks third in effectiveness according to the Prompt Report, ahead of zero-shot and ensembling. It works best when combined with specific evaluation criteria, giving the model concrete dimensions to assess rather than open-ended review.
Self-Consistency: Sampling Multiple Reasoning Paths
Self-consistency (Wang et al., 2022) extends chain-of-thought by generating multiple independent CoT reasoning paths for the same query, then selecting the most common answer through majority voting. The intuition: if a problem has a correct answer, multiple valid reasoning paths should converge on it, while errors in individual paths tend to be random and cancel out.
The procedure:
- Send the same CoT prompt to the model N times (typically N = 5-20) with temperature > 0.
- Extract the final answer from each response.
- Return the answer that appears most frequently.
Example: Self-consistency setup
[Prompt sent 5 times with temperature=0.7]
Q: A company's revenue grew 15% in Q1, declined 8% in Q2, and grew 22% in Q3.
If Q1 starting revenue was $1M, what is the revenue at the end of Q3?
Think step by step, then provide the final revenue figure.
Suppose the five responses produce final answers: $1.291M, $1.291M, $1.291M, $1.312M, $1.291M. The majority answer ($1.291M) is selected. The single outlier ($1.312M) likely contained an arithmetic error in one step.
Despite its theoretical appeal, empirical results from the Prompt Report indicate that self-consistency shows limited effectiveness gains compared to standard few-shot CoT. The accuracy improvement from zero-shot CoT (0.547) to zero-shot CoT with self-consistency (0.574) is modest, and the same pattern holds for few-shot variants (0.692 vs. 0.691). The computational cost of N parallel calls rarely justifies the marginal accuracy improvement for most production tasks.
Self-consistency is most valuable when: (a) the cost of a wrong answer is high, (b) the task involves numerical computation where small errors are common, and (c) latency constraints permit multiple parallel calls.
Tree of Thoughts: Hierarchical Exploration of Solution Spaces
Tree of Thoughts (ToT), introduced by Yao et al. (2023), generalizes chain-of-thought from a single linear reasoning path into a tree structure. At each step, the model generates multiple candidate "thoughts" (partial solutions), evaluates them, and selectively expands the most promising branches. Unpromising branches are pruned.
ToT operates through three components:
- Thought generation: At each node, produce K candidate next-steps.
- Evaluation: Score each candidate for correctness, relevance, or progress toward the goal.
- Search: Use BFS (breadth-first search) or DFS (depth-first search) to navigate the tree, backtracking from dead ends.
Example: ToT for a planning problem
Problem: Arrange a 3-day conference schedule with 12 speakers across 3 rooms,
subject to these constraints:
- No speaker conflicts (same speaker cannot be in two rooms simultaneously)
- Keynotes must be in the main hall
- Related topics should be in the same room
- Each room has a maximum capacity that varies
Step 1 - Generate 3 candidate day-1 schedules:
Candidate A: [Group speakers by topic cluster, assign keynotes to main hall first]
Candidate B: [Assign by speaker availability, fill rooms sequentially]
Candidate C: [Optimize for audience flow, minimize room transitions]
Step 2 - Evaluate each candidate:
- Candidate A: Satisfies topic clustering and keynote constraints. Score: 8/10.
- Candidate B: Creates speaker conflicts on day 2. Score: 3/10.
- Candidate C: Good flow but violates capacity constraint in room 2. Score: 6/10.
Step 3 - Expand Candidate A, prune Candidate B, attempt to fix Candidate C...
ToT is appropriate for combinatorial problems, planning tasks, puzzle-solving, and any scenario where the solution requires exploring and comparing alternative paths. The overhead is substantial: each node expansion requires a model call, and the tree can grow exponentially. For simple reasoning tasks, CoT is sufficient and far more efficient.
In practice, ToT can be implemented programmatically by wrapping model calls in a search algorithm. The model acts as both the generator (proposing next steps) and the evaluator (scoring candidates), while external code manages the tree traversal and pruning logic.
ReAct: Interleaving Reasoning with Action
ReAct (Yao et al., 2022) combines chain-of-thought reasoning with the ability to take external actions, such as searching a database, calling an API, or reading a document. The model alternates between "Thought" steps (internal reasoning) and "Action" steps (external operations), using the observations from actions to inform subsequent reasoning.
The ReAct loop:
- Thought: Reason about what information is needed.
- Action: Execute an external operation (search, lookup, calculate).
- Observation: Receive the result of the action.
- Repeat until the task is complete.
Example: ReAct for fact-based question answering
Question: What is the population difference between the two largest cities in Portugal?
Thought 1: I need to find the two largest cities in Portugal by population.
Action 1: Search["largest cities in Portugal by population"]
Observation 1: Lisbon (545,923), Porto (237,591), Vila Nova de Gaia (186,502)...
Thought 2: The two largest cities are Lisbon and Porto. I need to calculate the difference.
Action 2: Calculate[545923 - 237591]
Observation 2: 308,332
Thought 3: The population difference between Lisbon and Porto is 308,332.
Answer: The population difference between Lisbon and Porto is 308,332.
ReAct prevents hallucination by grounding reasoning in retrieved facts rather than parametric memory alone. It is the foundation pattern for agentic systems, where models interact with tools, databases, and external services.
The key design decision in ReAct is defining the action space: what tools are available, what their input/output formats are, and how the model selects among them. A well-defined action space constrains the model's behavior and reduces errors. A poorly defined one leads to tool misuse, infinite loops, or irrelevant actions.
ReAct is best suited for fact-based QA, data retrieval tasks, multi-step research, and any workflow where the model needs access to information not contained in its training data or context window.
ART: Automatic Reasoning and Tool-Use
ART (Automatic Reasoning and Tool-use) extends ReAct by automating the decomposition of tasks into subtasks and the selection of appropriate tools for each subtask. Where ReAct requires the user to define the reasoning/action structure, ART generates this structure automatically from a library of task templates and tool descriptions.
ART operates as follows:
- Receive a new task.
- Retrieve relevant task demonstrations from a library (similar to few-shot, but with tool-use examples).
- Decompose the task into subtasks, selecting tools for each.
- Execute the plan, pausing for tool outputs before continuing.
Example: ART for data analysis
Task: Analyze Q3 sales data and identify the top 3 underperforming regions.
[ART automatically decomposes into:]
Subtask 1: Retrieve Q3 sales data
Tool: database_query("SELECT region, revenue, target FROM sales WHERE quarter='Q3'")
Subtask 2: Calculate performance ratio per region
Tool: calculate(revenue / target for each region)
Subtask 3: Rank regions by performance ratio
Tool: sort(results, key=performance_ratio, order=ascending)
Subtask 4: Format top 3 underperforming regions with context
Tool: generate_report(bottom_3_regions, include=[gap_analysis, trend_comparison])
ART reduces the prompt engineering burden by shifting the task decomposition from the user to the model. It is particularly useful when building general-purpose agentic systems that must handle a wide variety of task types without task-specific prompt templates.
The tradeoff is reliability: automated decomposition can produce suboptimal plans, especially for novel tasks that do not closely match templates in the demonstration library. For high-stakes applications, a human-designed ReAct structure is more predictable than ART's automated approach.
Context Engineering: Beyond Prompt Engineering
Context engineering extends prompt engineering from "what you say" to "everything the model sees." While prompt engineering focuses on the instruction text itself, context engineering is the discipline of building a rich, comprehensive informational environment for the model, including examples, memory, state/history, retrieval results, structured outputs, and tool definitions.
The quality of this context is a primary factor in enabling advanced agentic performance. As the complexity of tasks grows, the model needs more than a well-crafted prompt; it needs the right information available at the right time.
Context Window Optimization
Three strategies keep context effective within token limits:
- Token budgeting: Optimize content density to stay within limits. Prune irrelevant information and balance cost vs. performance.
- Strategic truncation: Remove non-essential context and prioritize crucial information for each step or task.
- Context layering: Use caching and modular retrieval to structure long (and reusable) context windows efficiently.
Memory, Retrieval, and Modular Agents
Production systems often combine multiple context sources:
- Persistent memory systems carry forward relevant history, recursively updating internal state. These are proven to boost long-horizon agent reliability and efficiency.
- RAG modules provide LLMs with just-in-time facts and documents, grounding answers and reducing hallucination. Research shows RAG succeeds when context is not just "relevant" but "sufficient for solution," making re-rankers and sufficiency checks important for retrieval accuracy.
- Structured formats (Markdown, JSON, YAML, explicit delimiters) make context easier for models to parse and reason with. Designing context using explicit schemas, protocols, and modular blocks improves tool integration.
Practical Frameworks for Context Engineering
As systems grow in complexity, context engineering scales through several patterns:
- Context Engineering: Scale from prompts to examples to memory to modular agents to continuous context.
- Field Orchestration: Coordinate modules and retrievals in a unified pipeline.
- Audit and Prune: Trim irrelevant context, track residue, manage cost and reliability.
- Cognitive Recipes: Modularize reasoning steps for reuse and clarity.
The core principle is iterative: start small, rigorously add only missing elements, and measure impact on latency, effectiveness, and auditability at each stage.
Prompt Engineering Rules That Work
Beyond choosing the right technique, the quality of the prompt itself determines much of the outcome. Four rules consistently improve results across all techniques:
Be Clear and Direct
Use simple language and state what you want explicitly. Lead with a direct statement of the model's task. Use instructions and action verbs ("Write", "Create", "Generate") rather than questions. Avoid vague or roundabout phrasing.
Be Specific
Always use output guidelines. For troubleshooting, decision-making, critical thinking, or any situation where multiple angles should be considered, provide step-by-step instructions for the model to follow.
Provide Structure with XML Tags
When prompts contain multiple inputs, context blocks, or instructions, wrap each section in descriptive XML tags. This is most useful when including large amounts of context, mixing different types of content, or working with complex prompts that interpolate multiple variables. XML tags serve as clear delimiters that help the model distinguish instruction from data.
Provide Examples with Structure
Always use XML tags to structure examples clearly. Be explicit about what you are showing. Include examples that address your most common failure cases. Explain why your example outputs are considered ideal. Keep examples relevant to your specific task.
Prompting Myths to Avoid
Two common practices perform worse than expected in controlled experiments:
Role prompting ("You're a math professor, now look at this data set") shows measurable benefit only for creative writing tasks. For accuracy, reasoning, and factual tasks, it provides no improvement. This contradicts years of common practice but is consistently supported by the Prompt Report's empirical data.
Emotion prompting ("My career will be over if you get this wrong!") does not improve model performance. The model does not respond to emotional pressure, so these additions waste tokens without benefit.
A Decision Framework: Choosing the Right Technique
Selecting the right prompting technique requires evaluating four dimensions of the task: complexity, accuracy requirements, available resources (examples, tools, compute budget), and latency constraints.
Start simple. Add complexity only when justified by measurable improvement. This principle, validated across production implementations (including real-world deployments at Shopify), prevents the common mistake of reaching for advanced techniques before establishing a baseline. Most teams jump to fine-tuning because it sounds sophisticated. Smart teams start with prompts and prove value first.
The decision flow:
-
Try zero-shot first. If the task is simple classification, extraction, or generation with clear instructions, zero-shot may be sufficient. Test on a representative sample.
-
Add few-shot examples if zero-shot underperforms. This is the single highest-impact improvement for most tasks. Select 2-5 diverse, high-quality examples that cover edge cases.
-
Add chain-of-thought if reasoning is required. Multi-step calculations, logical inference, data analysis, and troubleshooting all benefit from CoT. The few-shot CoT combination (examples with reasoning chains) is the empirically strongest general-purpose technique.
-
Consider self-consistency for high-stakes numerical tasks. If a wrong answer has significant cost and the task involves computation, sampling 5-10 CoT paths with majority voting adds a safety margin.
-
Use Tree of Thoughts for combinatorial or planning problems. Scheduling, optimization, puzzle-solving, and multi-constraint satisfaction problems benefit from exploring and pruning alternative solution paths.
-
Use ReAct when external data is required. If the answer depends on information not in the model's training data or context, interleave reasoning with search, database queries, or API calls.
-
Use ART when building general-purpose agents. If the system must handle varied task types with a shared tool library, ART's automatic decomposition reduces per-task engineering effort.
-
Consider RAG independently. If external or changing data is involved, add retrieval augmentation at any tier. RAG is orthogonal to the prompting technique; you can combine RAG with zero-shot, few-shot, CoT, or ReAct.
-
Evaluate fine-tuning only after prompting plateaus. If you cannot achieve 80% accuracy with prompt engineering and you have 1,000+ high-quality examples, assess the ROI of fine-tuning. Otherwise, collect data while using prompts and iterate.
A critical insight from production deployments: most teams overinvest in technique sophistication and underinvest in prompt clarity. A well-written zero-shot prompt with clear instructions, specific output format requirements, and relevant context often outperforms a poorly constructed chain-of-thought prompt. Before escalating techniques, verify that the base prompt follows fundamental principles: be clear, be direct, be specific, provide structure (XML tags for multi-part inputs), and define the expected output format.
Quick Reference: Technique Selection by Task Type
| Task Type | Recommended Starting Technique | Escalation Path |
|---|---|---|
| Text classification | Zero-shot | Few-shot if edge cases arise |
| Data extraction | Zero-shot with output schema | Few-shot for complex formats |
| Arithmetic / logic | Zero-shot CoT | Few-shot CoT, then self-consistency |
| Code generation | Few-shot with examples | Few-shot CoT for complex algorithms |
| Multi-step analysis | Few-shot CoT | Self-consistency for validation |
| Planning / scheduling | CoT | Tree of Thoughts for constraint satisfaction |
| Fact-based QA | Zero-shot + RAG | ReAct for multi-hop questions |
| General-purpose agents | ReAct | ART for diverse task libraries |
Effectiveness Rankings
Empirical data from the Prompt Report, the most comprehensive systematic survey of prompting techniques (1,565 papers analyzed), ranks technique categories by effectiveness:
- Few-shot prompting (highest)
- Decomposition (break problems into subproblems)
- Self-criticism (critique and refine)
- Context/Information provision
- Ensembling (combining outputs from multiple prompts)
- Chain-of-thought (standalone, without few-shot)
- Role prompting (lowest for accuracy tasks)
Threats and bribes ("You must get this right or else") also rank at the bottom, confirming that emotional manipulation of models is ineffective.
Practical Recommendations
Iterate on prompts systematically. Define a goal, write an initial prompt, evaluate against a test set, apply a technique, and re-evaluate. This loop, repeated until performance stabilizes, is the core workflow of production prompt engineering.
Use XML tags for structure. When prompts contain multiple inputs, context blocks, or instructions, wrap each section in descriptive XML tags. This helps the model parse complex prompts and reduces confusion between instruction and data.
Provide output constraints. Specify the exact format, length, and content requirements of the expected response. Unconstrained generation produces varied formats that are harder to parse programmatically.
Version control your prompts. Treat prompts as code artifacts. Track changes, measure performance per version, and roll back when regressions occur. Prompt behavior can change across model versions, so a regression testing pipeline is essential for production systems.
Invest in context engineering as systems scale. As applications grow beyond single-prompt interactions into multi-step agentic workflows, the context surrounding each model call matters as much as the prompt itself. Budget tokens carefully, maintain persistent memory where needed, and structure all context with explicit schemas.
The gap between amateur and professional prompt engineering is not knowledge of advanced techniques. It is the discipline of systematic evaluation, the willingness to start simple, and the rigor to add complexity only when data justifies it.