A/B testShowing two versions to similar groups of users at the same time and comparing a metric, so you know which change caused the difference.
Acceptance criteriaConditions a piece of work must meet to count as done, often written as given, when, then.
AgentAn AI system that plans its own steps and uses tools to complete a goal.
APIA defined way for one program to ask another for data or actions, such as sending a prompt to a model.
BacklogThe ordered list of work a team might do, with the most valuable items at the top.
Balancing loopA feedback loop that pushes a system toward a limit or goal.
BiasSystematic unfairness in outputs, often learned from historical data.
Context engineeringChoosing what instructions, documents, history and tools go into a model's context.
Context windowThe maximum text a model can consider in one request.
Cycle timeHow long one item takes from start to finish.
Data driftReal-world data changing after a model was built, which lowers its accuracy.
Definition of DoneA team's shared checklist for when work is complete.
Double DiamondA design model with two phases of diverging and converging: one for the problem, one for the solution.
DPDPIndia's Digital Personal Data Protection Act, 2023: rules for how apps collect, use and protect personal data, built on clear consent.
EmbeddingA list of numbers representing meaning, used to find similar text.
Error analysisReading real outputs, noting failures and grouping them to decide what to fix.
EvalA repeatable test of an AI system's output quality.
Few-shot promptingGiving a model a few examples of the input and output you want.
Fine-tuningFurther training a model on specific examples to change its behaviour.
Golden setA fixed set of real inputs with known good answers, used to score every new version of an AI feature the same way. Start with 50 to 100 cases, hard ones included, and grow it from real failures.
GroundingGiving a model trusted source material so its answer rests on facts.
GuardrailA rule or check that blocks unsafe or off-topic inputs and outputs.
HallucinationFluent, confident output that is false or unsupported.
How Might WeAn open question that turns an insight into a design challenge.
Human in the loopA person reviews or approves AI output before it takes effect.
InferenceRunning a trained model to produce an output for a request.
INVESTGood user stories are Independent, Negotiable, Valuable, Estimable, Small and Testable.
Jobs to be doneThe progress a person is trying to make in a situation, which a product is hired to help with.
KanbanA method that visualizes work, limits work in progress and improves flow.
LatencyThe wait between a request and its response.
Leverage pointA place in a system where a small change causes a large effect.
LLMLarge language model: a model trained on huge amounts of text to predict and generate language.
LLM judgeUsing a model to score another model's output against a rubric.
MCPModel Context Protocol, an open standard for connecting AI models to tools and data.
Model cardA short document describing a model's purpose, data, performance and limits.
MVPMinimum viable product: the smallest thing that tests your riskiest assumption with real users.
North Star metricThe one metric that best captures the value users get from a product.
OKRObjectives and key results: a goal plus measurable outcomes that show progress.
PRDProduct requirements document: problem, users, scope, success metrics and out of scope.
PrecisionOf everything a system flagged, the share that was really right. High precision means few false alarms.
PromptThe instructions and content you give a model.
Prompt engineeringWriting, testing and refining the instructions, examples and context you give a model so it does a task well.
Prompt injectionText hidden in a document, web page or message that tries to override the model's instructions, for example to leak data.
RACIResponsible, Accountable, Consulted, Informed: who does what for a task or decision.
RAGRetrieval-augmented generation: fetching relevant documents and adding them to the prompt.
Reasoning modelA model that writes hidden working steps before it answers. It is often better on hard problems, but slower and costlier per task.
RecallOf everything that should have been flagged, the share the system caught. High recall means few misses.
Reinforcing loopA feedback loop where more leads to more, driving growth or collapse.
RetrospectiveA regular team meeting to inspect how work went and pick improvements.
RICEA prioritization score: reach times impact times confidence, divided by effort.
Risk registerA list of risks with likelihood, impact, owner and planned response.
ScrumAn agile framework of fixed-length sprints with set roles and events.
SprintA fixed time box, often two weeks, in which a team delivers a usable increment.
Stock and flowA stock is what builds up in a system. Flows are what add to it or drain it.
TemperatureA setting that controls how random a model's next-word choices are. Low gives steady, repeatable text. High gives more varied text.
TokenA chunk of text, often part of a word, that models read and write.
Tool useA model calling functions you define, like search, a calculator or an API.
Triple constraintThe trade-off between scope, time and cost in a project.
User storyA short description of a need: as a user, I want something, so that I get a benefit.
Vector databaseA store for embeddings that finds the passages closest in meaning to a question. RAG apps use one to retrieve context.
WIP limitA cap on how many items can be in progress at once.
WorkflowA fixed sequence of steps. In AI, a designed pipeline as opposed to a self-directed agent.