Prompt engineering course for companies
Prompt engineering is engineering, not copywriting
In production AI systems, the prompt determines whether a feature works or whether the helpdesk overflows. A good prompt engineer thinks in system prompts, few-shot examples, JSON schemas, retrieval context and evals, not in clever sentences for ChatGPT. Our course teaches technical teams how to design, version, test and monitor prompts as part of a serious software stack.
Discuss an in-house course View curriculumWho this course is for
This is not a "learn to write ChatGPT prompts" workshop. It is a technical training for teams who have LLMs in production or are about to deploy them, with all the versioning, evaluation and monitoring issues that come with that.
The course is designed for developers, data engineers, ML engineers, content teams working with structured generation, and customer service leads who steer LLM bots. Prerequisites: participants already work with an LLM API (OpenAI, Anthropic, Azure OpenAI or an open-source model via vLLM/Ollama), understand the difference between a chat completion and a completion endpoint, and have built at least one feature in which an LLM sits in a request-response cycle. We move beyond the basics: function calling, JSON mode, schema-constrained decoding, RAG pipelines and multi-agent handoffs are the standard.
What we explicitly do not do: share prompt tips you can find on LinkedIn, practise "act as a senior consultant" hacks, or claim that a good prompt solves all LLM problems. We treat prompt engineering as one layer in a larger AI system, alongside model choice, retriever tuning, guardrails and monitoring. If you only want to learn how to write a better ChatGPT prompt, this is the wrong place; other courses exist for that.
Curriculum: five modules, building on each other
The course runs from prompt architecture to agent orchestration. Each module combines theory with hands-on exercises based on participants' own use cases, not toy examples or invented tasks.
Prompt architecture
How to build a prompt for production: the difference between system and user prompts, role-based prompting, context management in long conversations, and when to truncate or summarise context. We cover prompt components (instructions, context, examples, schemas, output format) and how to turn them into a reusable template system so developers don't reinvent every prompt.
- System versus user prompt: what belongs where
- Role-based prompting and persona stability
- Context window management and token budgets
- Prompt templates and variable substitution without injection risk
Few-shot, zero-shot, chain-of-thought and ReAct
When should you choose which approach? Few-shot works well for classification and structured extraction, but costs tokens; chain-of-thought improves reasoning tasks but increases latency; ReAct is powerful for agent loops but unstable without strict schema control. We build all four and benchmark them against each other on a real task.
- Few-shot example selection and in-context learning
- Chain-of-thought (CoT) prompting and self-consistency
- ReAct: reasoning and acting in a single loop
- Trade-offs between accuracy, latency and cost
Function calling, tool use and structured output
A production LLM does not produce free text but a validated JSON object you can use downstream. We cover function calling (OpenAI, Anthropic), tool use with a tool router, JSON mode, and schema-constrained decoding via libraries such as Outlines or Instructor. Including retry logic for invalid outputs and how to prevent the model from emitting half-formed tool calls.
- Function calling specification and tool-routing patterns
- JSON mode versus schema-constrained decoding
- Pydantic and Zod schemas as the single source of truth
- Validation, retries and fallback strategies
Guardrails, jailbreak prevention and prompt injection
Any LLM system that processes user input is an attack surface. We cover prompt injection (direct and indirect, for example via retrieved RAG documents), jailbreak patterns and the well-known defences: input sanitisation, sandwich prompting, output filtering, and separating trusted and untrusted context. Including practical demos of current jailbreaks and how to block them without breaking your legitimate use case.
- Direct and indirect prompt injection
- Jailbreak categories and mitigation patterns
- Output filtering and content moderation
- Separating trusted and untrusted context
RAG prompting, agent prompting and multi-agent handoffs
Retrieval-augmented generation sounds simple — put documents in the context — but in production all the complexity lies in the prompting: how you instruct the model to use only cited facts, how you handle conflicting sources, and how you prevent retriever noise from confusing the model. We then build an agent loop with tool use and cover multi-agent handoffs: when agent A hands over to B, and how you make the shared state persistent.
- RAG prompts: citation enforcement and source grounding
- Integrating retriever output without hallucination
- Agent loops: planning, action, observation
- Multi-agent handoffs and shared memory
Evals: how you know a prompt works
A prompt that "feels good" during development is not a prompt that works in production. Evals are the unit tests of AI systems: without evals you cannot tell whether your new prompt version is better or worse than the previous one. This module covers eval frameworks and how to build them into your CI pipeline.
Promptfoo
Open-source CLI for prompt testing. Define test cases in YAML, run them against multiple models, and compare outputs side by side. Suitable for regression testing of prompt changes in CI. Works well for structured-output use cases where you can define an expected JSON per case.
Langfuse
Tracing and eval platform for LLM apps. Logs every call (prompt, response, latency, cost) and lets you run evals on production traces. Open source and self-hostable — important for teams with EU data residency requirements. Integrates with LangChain, LlamaIndex and raw OpenAI/Anthropic SDKs.
Helicone
Observability layer you place between your app and the LLM API. Gives insight per request into latency, cost, cache hits and error rates. Combines well with Promptfoo or Langfuse: Helicone for production telemetry, Promptfoo for offline regression.
Custom evals: LLM-as-judge
For open-ended outputs, a keyword match does not work. LLM-as-judge uses a strong model to score another model's output against rubrics (correctness, tone, format compliance). We cover judge bias, calibration, and how to prevent your judge from making the same mistakes as your production model.
Golden datasets
Eval quality is dataset quality. We teach participants how to build a golden dataset from real production interactions: edge cases, regression cases, adversarial inputs. Includes annotation workflows and how to detect dataset drift.
A/B testing in production
Offline evals don't cover everything. We cover feature flags for prompt versions, traffic splitting by user cohort, and how to measure a statistically significant prompt uplift without harming users if a variant performs poorly.
Production issues you won't learn anywhere else
A prompt that works on your laptop is not the same as a prompt that runs stably for six months. This module covers the operational reality: versioning, drift, costs and caching.
Prompt versioning is not optional but mandatory. Every prompt that touches production gets a version number, a changelog and a rollback path, just like code. We cover prompt versioning in Git (prompts-as-code), in a prompt management tool (Langfuse, PromptLayer), and hybrid models where prompts live in a database with an audit trail.
Drift monitoring is the second part. A prompt that works today can perform worse in three months because the underlying model has been updated (silent model updates by API providers), because the input distribution has shifted, or because an upstream system has changed its output format. We cover drift detection through continuous evals on a fixed set, semantic similarity checks on outputs, and alerting on score degradation.
Cost optimisation is the third. Naive LLM implementations can be far more expensive than necessary. We cover prompt caching (Anthropic prompt caching, OpenAI's automatic caching, Redis-based response caching), model routing (a cheap model for simple queries, a strong model for complex ones), batch processing for asynchronous workloads, and how to enforce a cost budget per request without degrading your service during peak usage.
Finally, the practical things nobody ever writes about. How do you handle provider rate limits (exponential backoff, request queueing, multi-provider fallback)? How do you debug a prompt that fails "sometimes"? How do you roll out a prompt update without downtime? And how do you explain to compliance that you aren't generating made-up content but a grounded answer from your RAG pipeline?
Hands-on: participants work on their own use cases
The course is not a series of slides. Every module has hands-on assignments in which participants write, evaluate and revise prompts for their own production use cases. No abstract exercises: the prompt you build in the course can be deployed on Friday.
CRM AI: lead enrichment and classification
A common use case: the CRM system must automatically classify incoming leads, enrich them with public data and route them to the right sales rep. Participants build a function-calling prompt with a JSON schema, a retry loop for invalid outputs, and evals on a dataset of historical leads. Includes edge-case handling (incomplete data, multi-language input, ambiguous cases).
Customer service bot with RAG
A customer service bot that draws answers from a knowledge base: manuals, FAQs, policy documents. Participants design the RAG prompt (citation enforcement, fallback when there is no match, escalation to a human), build an eval set with edge cases, and test prompt-injection resistance with adversarial inputs.
Internal wiki RAG for engineering teams
A second RAG case: a Q&A system that searches internal documentation: Confluence, Notion, Google Docs. Different challenges from customer service: more technical jargon, code snippets in retrieval, multi-document answers, and ACL compliance (which user may see which documents). Participants build a prompt that draws exclusively from permitted sources.
Document extraction pipeline
Extract structured data from PDFs, contracts or forms using function calling and schema-constrained decoding. Participants define a Pydantic schema, write the extraction prompt, build retries for parse failures and measure extraction accuracy against a manually labelled set.
If you don't have your own use case, you'll be given a realistic case study, for example a document extraction pipeline built on publicly available contracts, or a RAG bot built on open-source documentation. Most teams, however, arrive with several production use cases they want to tackle during the course.
Tools and frameworks we cover
The course is provider-agnostic in its core concepts, but the hands-on work focuses on the tools most commonly used in production. Participants gain practical experience with multiple models and frameworks so they can make well-considered choices, rather than simply sticking with "whatever we happen to know".
Importantly, we teach participants why a tool suits a given use case, not how to name-drop every tool. LangChain is powerful but overkill for simple apps; LiteLLM is a thin abstraction for multi-provider routing; Instructor is excellent for schema-constrained extraction but adds a layer. Which combination fits your stack depends on team skills, deployment environment and data residency requirements, and we guide participants through exactly that.
Case walkthrough: from naive prompt to production-grade
A typical exercise from the course, in brief: how a quote-extraction prompt evolves from a 200-token zero-shot attempt into an instrumented pipeline with evals, caching and monitoring.
Iteration 1 — Naive zero-shot
A developer writes a prompt: "Extract the customer name, order ID and total amount from this email." It works on some of the inputs. On the rest, the model extracts incorrect fields, hallucinates IDs or refuses to answer multilingual emails. There's no schema, no retries and no evals.
Iteration 2 — Few-shot with JSON schema
Add a few few-shot examples covering variety (English, Dutch, with VAT, without VAT), and function calling with a Pydantic schema that defines customer_name, order_id and total_amount_eur strictly typed fields. Accuracy improves noticeably, but costs double because of the extra context tokens, and with unusual formats the model still produces invalid JSON.
Iteration 3 — Schema-constrained decoding + retry
Replace this with schema-constrained decoding via Instructor, so the model can no longer structurally produce invalid JSON. Add retry logic with exponential backoff for parse failures. Set up an eval set in Promptfoo to make every prompt change measurable. Latency barely increases.
Iteration 4 — Production instrumentation
Langfuse logs every call with prompt version, cost and latency. Helicone caches similar requests (identical email templates yield cache hits). An A/B test in production compares two prompt versions (an extra disambiguation instruction for amounts in currencies other than EUR). A drift eval runs each night on new production traces to detect regressions.
What you learn from this flow
- Never start with a complex prompt; measure the baseline first.
- Schema-constrained decoding solves a whole class of problems that retries alone cannot fix.
- Evals are the blocker: without an eval set, you can't tell iteration 1 apart from iteration 4.
- Caching is free cost saving as soon as you call the same model several times with similar input.
- A prompt is never "finished"; drift monitoring catches the regressions you would never see in development.
Frequently asked questions about the prompt engineering course
Planning a prompt engineering course for your team?
Tell us which LLM features you have in production or want to build, what team level you are starting from, and which deployment environment you use. We will put together a curriculum that fits your stack and use cases: no template, no generic slides.
Book an intake call