🧭 Prompt Engineering — Engineering Reference

A code-first, production-grade prompt-engineering curriculum for mid-to-senior developers moving toward staff/principal roles. Each chapter is structured around annotated code blocks — complex implementations, anti-patterns with fixes, performance tricks, and edge-case failure modes — rather than prose-heavy tutorials. Every technique is shown as it's used in real production systems, with dense inline comments explaining the underlying mechanism.

How to Use This Reference

  1. Read sequentially (01 → 20) for a structured path from LLM mechanics to production evaluation.
  2. Jump to a chapter as a reference when you hit a specific prompting problem in the wild — each is self-contained.
  3. Run the exercises in chapter 20 after every few chapters, not just at the end.
  4. Treat every code example as a starting point — adapt the patterns to your domain, test against your own eval set (Chapter 19), and verify behavior on YOUR target model (Chapter 16).

Prerequisites

  • Familiarity with at least one LLM API (Anthropic, OpenAI, or equivalent) — you've made API calls and understand the request/response shape.
  • Working knowledge of Python — code examples use Python with anthropic SDK patterns, but the concepts transfer to any language.
  • Understanding of basic software engineering practices (versioning, testing, CI) — this reference treats prompts as production code, not creative writing.

Curriculum

Part I — Foundations

#TopicWhy It Matters
01Introduction & How LLMs WorkNext-token prediction, tokenization, context budgets — the mechanistic model every technique builds on.
02Anatomy of a PromptRole hierarchy, instruction/context/data separation, delimiter strategy — shown as real API request bodies.
03Zero-Shot & Few-Shot PromptingExample selection, ordering effects, diminishing returns, and the accidental-pattern trap.
04Clarity & SpecificityEliminating ambiguity through checkable constraints, positive framing, and anti-rebound instruction design.

Part II — Core Techniques

#TopicWhy It Matters
05Chain-of-Thought PromptingCoT as self-generated context, extended thinking, and the fluent-but-wrong failure mode.
06Role & Persona PromptingConditioning signals, sycophancy mitigation, anti-caving instructions, and persona drift.
07Output Formatting & Structured DataPrompted vs. API-enforced schemas, function calling as structured output, and parsing failure modes.
08Context & Memory ManagementSliding window, summarization, structured memory, prompt caching, and the layered production architecture.
09Iterative Refinement & Prompt TestingEval sets, A/B testing, LLM-as-judge, and the iteration loop that separates engineering from tweaking.

Part III — Advanced Techniques

#TopicWhy It Matters
10Decomposition & Task BreakdownPipelines of focused prompts, structured handoffs, parallel vs. sequential, and error compounding.
11Self-Consistency & VerificationSampling multiple paths, majority voting, external ground-truth checks, and multi-model cross-checking.
12Retrieval-Augmented Generation (RAG)Grounding instructions, citation verification, contradiction handling, and chunk placement strategy.
13Tool Use & Function CallingTool definitions, the calling loop, result formatting, error handling, and the tool-vs-prompt decision.
14Multi-Agent & Agentic WorkflowsOrchestrator/sub-agent patterns, structured handoffs, synthesis design, and human-in-the-loop enforcement.

Part IV — Model-Specific & Practical Craft

#TopicWhy It Matters
15Working with ClaudeXML tags, system prompt structure, extended thinking, literal instruction-following, and pushback encouragement.
16Working with GPT & Other ModelsPortability, OpenAI conventions, reasoning-optimized models, open-weight chat templates, and graceful degradation.
17Handling Hallucination & UncertaintyCalibrated uncertainty, grounding with citations, explicit I-don't-know permission, and domain-specific risk patterns.

Part V — Production & Safety

#TopicWhy It Matters
18Prompt Injection & SecurityDirect and indirect injection, defense-in-depth, architectural safeguards, and the SQL-injection analogy.
19Evaluating & Testing Prompts at ScaleEval harnesses, LLM-as-judge calibration, CI-gated regression testing, cost/latency metrics, and statistical significance.
20Exercises & Project IdeasFrom beginner drills to a self-hosted red-team bounty — where the curriculum turns into judgment.

Learning Path Suggestions

If you're a developer building LLM features into a product

Read 01–09 in order — don't skip the foundations even if you're experienced with APIs, since most production prompt bugs trace back to a Part I or II concept applied sloppily. Read 12, 13, and 19 closely. Read 15 or 16 depending on which model you're shipping with. Treat chapter 19's eval-harness pattern as non-optional before shipping to real users.

If you're building agents or tool-using systems

Skim 01–09. Read 10, 11, 13, and 14 carefully — this is the core of agentic design. Read 17 before you trust any agent's intermediate claims. Read 18 before you give an agent access to anything that matters. Finish with exercises 9, 12, and 13 in chapter 20.

If your focus is safety, security, or red-teaming

Read 01–04 for the mental model, then jump straight to 17 and 18. Read 19 to understand how injection resistance gets regression-tested rather than checked once. Do exercises 10 and 13 in chapter 20, and treat project 14 (the self-hosted prompt injection bug bounty) as the capstone.

Companion Resources