Prompt engineering course for companies

Prompt engineering is engineering, not copywriting

In production AI systems, the prompt determines whether a feature works or whether the helpdesk overflows. A good prompt engineer thinks in system prompts, few-shot examples, JSON schemas, retrieval context and evals, not in clever sentences for ChatGPT. Our course teaches technical teams how to design, version, test and monitor prompts as part of a serious software stack.

System prompt architecture Function calling RAG prompting Evals & LLM-as-judge Prompt injection defence
Discuss an in-house course View curriculum
{ "ok": true }

Who this course is for

This is not a "learn to write ChatGPT prompts" workshop. It is a technical training for teams who have LLMs in production or are about to deploy them, with all the versioning, evaluation and monitoring issues that come with that.

The course is designed for developers, data engineers, ML engineers, content teams working with structured generation, and customer service leads who steer LLM bots. Prerequisites: participants already work with an LLM API (OpenAI, Anthropic, Azure OpenAI or an open-source model via vLLM/Ollama), understand the difference between a chat completion and a completion endpoint, and have built at least one feature in which an LLM sits in a request-response cycle. We move beyond the basics: function calling, JSON mode, schema-constrained decoding, RAG pipelines and multi-agent handoffs are the standard.

What we explicitly do not do: share prompt tips you can find on LinkedIn, practise "act as a senior consultant" hacks, or claim that a good prompt solves all LLM problems. We treat prompt engineering as one layer in a larger AI system, alongside model choice, retriever tuning, guardrails and monitoring. If you only want to learn how to write a better ChatGPT prompt, this is the wrong place; other courses exist for that.

Curriculum: five modules, building on each other

The course runs from prompt architecture to agent orchestration. Each module combines theory with hands-on exercises based on participants' own use cases, not toy examples or invented tasks.

Prompt architecture

How to build a prompt for production: the difference between system and user prompts, role-based prompting, context management in long conversations, and when to truncate or summarise context. We cover prompt components (instructions, context, examples, schemas, output format) and how to turn them into a reusable template system so developers don't reinvent every prompt.

  • System versus user prompt: what belongs where
  • Role-based prompting and persona stability
  • Context window management and token budgets
  • Prompt templates and variable substitution without injection risk

Few-shot, zero-shot, chain-of-thought and ReAct

When should you choose which approach? Few-shot works well for classification and structured extraction, but costs tokens; chain-of-thought improves reasoning tasks but increases latency; ReAct is powerful for agent loops but unstable without strict schema control. We build all four and benchmark them against each other on a real task.

  • Few-shot example selection and in-context learning
  • Chain-of-thought (CoT) prompting and self-consistency
  • ReAct: reasoning and acting in a single loop
  • Trade-offs between accuracy, latency and cost

Function calling, tool use and structured output

A production LLM does not produce free text but a validated JSON object you can use downstream. We cover function calling (OpenAI, Anthropic), tool use with a tool router, JSON mode, and schema-constrained decoding via libraries such as Outlines or Instructor. Including retry logic for invalid outputs and how to prevent the model from emitting half-formed tool calls.

  • Function calling specification and tool-routing patterns
  • JSON mode versus schema-constrained decoding
  • Pydantic and Zod schemas as the single source of truth
  • Validation, retries and fallback strategies

Guardrails, jailbreak prevention and prompt injection

Any LLM system that processes user input is an attack surface. We cover prompt injection (direct and indirect, for example via retrieved RAG documents), jailbreak patterns and the well-known defences: input sanitisation, sandwich prompting, output filtering, and separating trusted and untrusted context. Including practical demos of current jailbreaks and how to block them without breaking your legitimate use case.

  • Direct and indirect prompt injection
  • Jailbreak categories and mitigation patterns
  • Output filtering and content moderation
  • Separating trusted and untrusted context

RAG prompting, agent prompting and multi-agent handoffs

Retrieval-augmented generation sounds simple — put documents in the context — but in production all the complexity lies in the prompting: how you instruct the model to use only cited facts, how you handle conflicting sources, and how you prevent retriever noise from confusing the model. We then build an agent loop with tool use and cover multi-agent handoffs: when agent A hands over to B, and how you make the shared state persistent.

  • RAG prompts: citation enforcement and source grounding
  • Integrating retriever output without hallucination
  • Agent loops: planning, action, observation
  • Multi-agent handoffs and shared memory

Evals: how you know a prompt works

A prompt that "feels good" during development is not a prompt that works in production. Evals are the unit tests of AI systems: without evals you cannot tell whether your new prompt version is better or worse than the previous one. This module covers eval frameworks and how to build them into your CI pipeline.

Promptfoo

Open-source CLI for prompt testing. Define test cases in YAML, run them against multiple models, and compare outputs side by side. Suitable for regression testing of prompt changes in CI. Works well for structured-output use cases where you can define an expected JSON per case.

Langfuse

Tracing and eval platform for LLM apps. Logs every call (prompt, response, latency, cost) and lets you run evals on production traces. Open source and self-hostable — important for teams with EU data residency requirements. Integrates with LangChain, LlamaIndex and raw OpenAI/Anthropic SDKs.

Helicone

Observability layer you place between your app and the LLM API. Gives insight per request into latency, cost, cache hits and error rates. Combines well with Promptfoo or Langfuse: Helicone for production telemetry, Promptfoo for offline regression.

Custom evals: LLM-as-judge

For open-ended outputs, a keyword match does not work. LLM-as-judge uses a strong model to score another model's output against rubrics (correctness, tone, format compliance). We cover judge bias, calibration, and how to prevent your judge from making the same mistakes as your production model.

Golden datasets

Eval quality is dataset quality. We teach participants how to build a golden dataset from real production interactions: edge cases, regression cases, adversarial inputs. Includes annotation workflows and how to detect dataset drift.

A/B testing in production

Offline evals don't cover everything. We cover feature flags for prompt versions, traffic splitting by user cohort, and how to measure a statistically significant prompt uplift without harming users if a variant performs poorly.

Production issues you won't learn anywhere else

A prompt that works on your laptop is not the same as a prompt that runs stably for six months. This module covers the operational reality: versioning, drift, costs and caching.

Prompt versioning is not optional but mandatory. Every prompt that touches production gets a version number, a changelog and a rollback path, just like code. We cover prompt versioning in Git (prompts-as-code), in a prompt management tool (Langfuse, PromptLayer), and hybrid models where prompts live in a database with an audit trail.

Drift monitoring is the second part. A prompt that works today can perform worse in three months because the underlying model has been updated (silent model updates by API providers), because the input distribution has shifted, or because an upstream system has changed its output format. We cover drift detection through continuous evals on a fixed set, semantic similarity checks on outputs, and alerting on score degradation.

Cost optimisation is the third. Naive LLM implementations can be far more expensive than necessary. We cover prompt caching (Anthropic prompt caching, OpenAI's automatic caching, Redis-based response caching), model routing (a cheap model for simple queries, a strong model for complex ones), batch processing for asynchronous workloads, and how to enforce a cost budget per request without degrading your service during peak usage.

Finally, the practical things nobody ever writes about. How do you handle provider rate limits (exponential backoff, request queueing, multi-provider fallback)? How do you debug a prompt that fails "sometimes"? How do you roll out a prompt update without downtime? And how do you explain to compliance that you aren't generating made-up content but a grounded answer from your RAG pipeline?

Hands-on: participants work on their own use cases

The course is not a series of slides. Every module has hands-on assignments in which participants write, evaluate and revise prompts for their own production use cases. No abstract exercises: the prompt you build in the course can be deployed on Friday.

CRM AI: lead enrichment and classification

A common use case: the CRM system must automatically classify incoming leads, enrich them with public data and route them to the right sales rep. Participants build a function-calling prompt with a JSON schema, a retry loop for invalid outputs, and evals on a dataset of historical leads. Includes edge-case handling (incomplete data, multi-language input, ambiguous cases).

Customer service bot with RAG

A customer service bot that draws answers from a knowledge base: manuals, FAQs, policy documents. Participants design the RAG prompt (citation enforcement, fallback when there is no match, escalation to a human), build an eval set with edge cases, and test prompt-injection resistance with adversarial inputs.

Internal wiki RAG for engineering teams

A second RAG case: a Q&A system that searches internal documentation: Confluence, Notion, Google Docs. Different challenges from customer service: more technical jargon, code snippets in retrieval, multi-document answers, and ACL compliance (which user may see which documents). Participants build a prompt that draws exclusively from permitted sources.

Document extraction pipeline

Extract structured data from PDFs, contracts or forms using function calling and schema-constrained decoding. Participants define a Pydantic schema, write the extraction prompt, build retries for parse failures and measure extraction accuracy against a manually labelled set.

If you don't have your own use case, you'll be given a realistic case study, for example a document extraction pipeline built on publicly available contracts, or a RAG bot built on open-source documentation. Most teams, however, arrive with several production use cases they want to tackle during the course.

Tools and frameworks we cover

The course is provider-agnostic in its core concepts, but the hands-on work focuses on the tools most commonly used in production. Participants gain practical experience with multiple models and frameworks so they can make well-considered choices, rather than simply sticking with "whatever we happen to know".

OpenAI GPT-4o / GPT-4-turbo Anthropic Claude Azure OpenAI Google Gemini Mistral Ollama / vLLM (local) LangChain LlamaIndex LiteLLM Instructor Outlines Pydantic Promptfoo Langfuse Helicone PromptLayer Weights & Biases Weave Pinecone / Qdrant / pgvector

Importantly, we teach participants why a tool suits a given use case, not how to name-drop every tool. LangChain is powerful but overkill for simple apps; LiteLLM is a thin abstraction for multi-provider routing; Instructor is excellent for schema-constrained extraction but adds a layer. Which combination fits your stack depends on team skills, deployment environment and data residency requirements, and we guide participants through exactly that.

Case walkthrough: from naive prompt to production-grade

A typical exercise from the course, in brief: how a quote-extraction prompt evolves from a 200-token zero-shot attempt into an instrumented pipeline with evals, caching and monitoring.

Iteration 1 — Naive zero-shot

A developer writes a prompt: "Extract the customer name, order ID and total amount from this email." It works on some of the inputs. On the rest, the model extracts incorrect fields, hallucinates IDs or refuses to answer multilingual emails. There's no schema, no retries and no evals.

Iteration 2 — Few-shot with JSON schema

Add a few few-shot examples covering variety (English, Dutch, with VAT, without VAT), and function calling with a Pydantic schema that defines customer_name, order_id and total_amount_eur strictly typed fields. Accuracy improves noticeably, but costs double because of the extra context tokens, and with unusual formats the model still produces invalid JSON.

Iteration 3 — Schema-constrained decoding + retry

Replace this with schema-constrained decoding via Instructor, so the model can no longer structurally produce invalid JSON. Add retry logic with exponential backoff for parse failures. Set up an eval set in Promptfoo to make every prompt change measurable. Latency barely increases.

Iteration 4 — Production instrumentation

Langfuse logs every call with prompt version, cost and latency. Helicone caches similar requests (identical email templates yield cache hits). An A/B test in production compares two prompt versions (an extra disambiguation instruction for amounts in currencies other than EUR). A drift eval runs each night on new production traces to detect regressions.

What you learn from this flow

  1. Never start with a complex prompt; measure the baseline first.
  2. Schema-constrained decoding solves a whole class of problems that retries alone cannot fix.
  3. Evals are the blocker: without an eval set, you can't tell iteration 1 apart from iteration 4.
  4. Caching is free cost saving as soon as you call the same model several times with similar input.
  5. A prompt is never "finished"; drift monitoring catches the regressions you would never see in development.

Frequently asked questions about the prompt engineering course

What prior knowledge do participants need?
Participants should have built at least one feature that calls an LLM via API, understand the difference between system and user prompts, and be fluent in Python or TypeScript. This course is not for beginners. If you are only just learning what a prompt is, we recommend a general introduction to LLMs first. We take function calling, JSON mode and RAG as given from day one.
What is the difference between prompt engineering and fine-tuning?
Prompt engineering adjusts the behaviour of a general-purpose model through instructions, examples and context, without changing the model weights. Fine-tuning adjusts the weights using a training dataset. For most production use cases, prompt engineering combined with RAG is sufficient. Fine-tuning becomes relevant only for specific style transfer, high-volume cost optimisation, or tasks that the base prompt architecture cannot solve. We cover when each tool is the right choice.
Do you work with OpenAI, Anthropic or open-source models?
Several. The concepts are provider-agnostic, but the syntax differs: function calling with OpenAI is not the same as tool use with Anthropic, and open-source models served via vLLM or Ollama have their own quirks. Participants gain experience with at least two providers and complete at least one exercise on a locally hosted model for data-residency cases.
How do you handle EU residency and privacy requirements?
We cover several deployment models: API providers with EU data centres (Azure OpenAI EU, Anthropic via AWS Bedrock EU), self-hosted open-source models via vLLM or Ollama, and hybrid architectures where sensitive extraction happens locally and only anonymised context is sent to an external API. Which model fits depends on the legal context of the client.
How long does the course take and how is it delivered?
The standard format is several days spread over two to three weeks: one day of theory and concepts, followed by hands-on days with your own use cases. Between sessions, participants work on assignments in their own codebase. We deliver on-site in-house, online via remote pairing, or in a hybrid format. We tune the exact schedule to the team composition and existing projects.
What determines the investment for an in-house course?
The investment depends on the number of participants, the degree of customisation to your own use cases, and whether we also work hands-on alongside the course on production prompts. A standard curriculum for a team differs from a course that also audits and restructures your existing prompt stack. We prepare a quote in advance based on scope.
Do you offer support after the course?
Yes, for teams that want it. Possible forms include a follow-up review a few weeks later, in which we assess the prompts teams have put into production since the course, an ongoing prompt engineering retainer, or a hands-on build assignment where we work alongside the team on the first production feature. The course itself is not a black-box finish but, where it fits, the start of a working relationship.
How does this course relate to an AI discovery workshop?
An AI discovery workshop is strategic: which AI use cases fit our business, what is the roadmap, and where is the business value. This prompt engineering course is technical and hands-on: how do you build the prompts used in those use cases. Many clients first run the discovery workshop with decision-makers and then the prompt engineering course with the implementation team.

Planning a prompt engineering course for your team?

Tell us which LLM features you have in production or want to build, what team level you are starting from, and which deployment environment you use. We will put together a curriculum that fits your stack and use cases: no template, no generic slides.

Book an intake call

Edit content