789 terms, 9 layers, 28 domains
Every word the field uses, explained once and properly — searchable, filterable by topic and level, and cross-linked so one term leads to the next.
A
- A* SearchThe classic best-first pathfinding algorithm combining path cost with a heuristic estimate — still everywhere from games to robotics.first seen 1968Advanced
- A/B TestingComparing two model or product variants on live traffic to measure which performs better.first seen 2000Advanced
- A2A (Agent2Agent)Google’s open protocol for agents from different vendors to discover each other and collaborate — a complement to MCP, which connects agents to tools.first seen 2025Advanced
- Abductive ReasoningInference to the best explanation — reasoning backwards from an observation to its most plausible cause; Peirce’s third mode next to deduction and induction.first seen 1878Expert
- Ablation StudyRemoving one component of a system at a time to measure how much it contributes to performance.first seen 1990Expert
- AbliterationFinding the refusal direction in a model’s residual stream and projecting it out of the weights — a targeted way to strip safety behaviour from open weights without retraining. Cheap, effective, and it measurably degrades the model elsewhere (see KL Divergence).first seen 2024Expert
- AccuracyThe fraction of predictions a model gets right — intuitive, but misleading on imbalanced data.first seen 1950Beginner
- Activation FunctionThe non-linearity (ReLU, GELU, sigmoid…) applied after each neuron that lets networks model complex functions.first seen 1943Advanced
- ActivationsThe intermediate numbers flowing through a network while it processes an input — not the weights, which stay fixed between runs. Interpretability, steering and abliteration all operate on activations.first seen 1986Advanced
- Active InferenceFriston’s free-energy framework: organisms (and agents) both perceive and act to minimize surprise about their world — an influential neuroscience lens on intelligence.first seen 2010Expert
- Active LearningA training loop where the model picks the most informative examples for humans to label next.first seen 1988Advanced
- Active ParametersThe share of an MoE model’s weights that actually fires for a single token — a 200B model with 20B active costs roughly 20B worth of compute per token but still needs all 200B in memory.first seen 2023Advanced
- Actor-CriticReinforcement learning setup pairing a policy (“actor”) with a value estimator (“critic”) that guides its updates.first seen 1983Expert
- Adam OptimizerThe most widely used neural-network optimizer; adapts each parameter’s learning rate using gradient statistics.first seen 2014Advanced
- Adapter (PEFT)Small trainable modules inserted into a frozen pretrained model so it can be adapted cheaply to new tasks.first seen 2019Advanced
- Adversarial ExampleAn input with tiny, often invisible perturbations crafted to make a model fail — a stop sign read as a speed limit.first seen 2013Advanced
- Affective ComputingSystems that recognize, interpret and simulate human emotion — founded as a field by Rosalind Picard at MIT.first seen 1995Advanced
- AgentAn AI system that perceives its environment and takes actions toward a goal — today usually an LLM that plans, uses tools and iterates with limited supervision.first seen 1959Beginner
- Agentic AISystems in which autonomy is the design principle — models that decide what to do next, run for hours, coordinate other agents and are supervised by exception rather than by prompt. «Agent» names the unit, «agentic» the degree: Gartner’s 2025 hype-cycle peak and, in the same breath, its warning that most «agentic» products are chatbots with a tool call.first seen 2024Beginner
- Agentic EngineeringBuilding software by directing coding agents — specs, tests, reviews and checkpoints instead of typing the code yourself. Andrej Karpathy, who coined «vibe coding» in February 2025, declared it passé a year later and named its disciplined successor: the engineer writes the goal and the guardrails, agents write and run the code, humans keep the oversight. Vendors sell the same shift under their own labels (AWS: Kiro, «frontier agents»); this is the neutral term.first seen 2026Beginner
- Agentic RAGRetrieval run by an agent loop — the model decides what to search, judges the results and searches again — instead of one fixed retrieve-then-answer pass.first seen 2024Advanced
- Agentic WorkflowA multi-step process in which a model plans, calls tools, checks results and retries until a task is done.first seen 2023Advanced
- AGI (Artificial General Intelligence)Hypothetical AI with human-level competence across virtually all cognitive tasks, not just narrow domains.first seen 1997Beginner
- AI AcceleratorUmbrella term for chips specialized for neural-network math — GPUs, TPUs, NPUs and their growing zoo of rivals.first seen 2016Advanced
- AI AlignmentThe research problem of making AI systems reliably pursue the goals and values their operators and society intend.first seen 2014Beginner
- AI Bill of Materials (AIBOM)An inventory of what went into an AI system — base model, fine-tuning data, libraries, licences, evaluation results — modelled on the software bill of materials. What auditors, EU AI Act compliance and enterprise procurement increasingly ask for.first seen 2023Advanced
- AI BubbleThe recurring worry that AI investment has outrun real value — the boom-side companion to the bust called AI Winter.first seen 2024Beginner
- AI Conferences (NeurIPS, ICML, ICLR, CVPR, ACL)Where the field publishes: NeurIPS (since 1987) and ICML for machine learning, ICLR for deep learning, CVPR for vision, ACL for language. Acceptance is the academic currency — and the papers appear on arXiv months before the meeting, so the conference is where the community meets, not where it learns the news.first seen 1987Advanced
- AI CopilotAn assistant that suggests, drafts and supports inside your workflow while you stay in control — popularized by GitHub Copilot.first seen 2021Beginner
- AI Data CenterCompute campuses purpose-built for training and serving models — dense GPU racks, exotic cooling and power draws that make utilities nervous.first seen 2022Beginner
- AI Energy UseData centres already draw around 1.5% of the world’s electricity and AI is the fastest-growing part — a single frontier training run consumes gigawatt-hours, a chat query several times the energy of a search. Switzerland’s data centres use about 4% of national electricity; the debate is over grid, cooling water and where the power comes from.first seen 2023Beginner
- AI EthicsThe field studying the moral questions of building and deploying AI — fairness, accountability, transparency, impact on people.first seen 2015Beginner
- AI for ScienceThe umbrella term for models that accelerate research itself — protein structure (AlphaFold, Nobel Prize 2024), weather, materials discovery (DeepMind’s GNoME proposed 2.2 million new crystals in 2023), fusion-plasma control, drug candidates. The claim is a decade of lab work per year; the caveat is that a prediction still has to be synthesised and measured before it counts.first seen 2020Beginner
- AI PsychosisMedia label (not a clinical diagnosis) for obsessive chatbot fixation feeding delusional beliefs — a flashpoint in the AI-companionship debate.first seen 2025Beginner
- AI SafetyThe field working to prevent harm from AI systems, from today’s failures to hypothetical loss of control — alignment is its core technical problem.first seen 2015Beginner
- AI Safety InstituteGovernment bodies (UK, US, EU…) that evaluate frontier models for dangerous capabilities before and after release.first seen 2023Beginner
- AI TOPSTera-operations per second at low precision — the headline speed number of NPUs and AI PCs; compare with care, since precision and sparsity tricks inflate it.first seen 2018Advanced
- AI Weather ModelForecast models learned from forty years of reanalysis data instead of solving the atmosphere’s equations — GraphCast (DeepMind), FourCastNet (NVIDIA), Pangu-Weather (Huawei), all 2022–23 — that beat the physical models on most ten-day scores in seconds on a single GPU. Since 2025 ECMWF runs its own AIFS operationally next to the classic forecast; MeteoSwiss runs its numerical model on Alps.first seen 2023Beginner
- AI WinterA period of collapsed funding and interest after hype outruns results — the field froze in the mid-1970s and late 1980s.first seen 1984Beginner
- AI WormA self-replicating prompt that spreads from one AI assistant to the next — the Morris II demo (2024) hid it in an e-mail so that each mail agent reading it forwarded it and leaked data. Theoretical for now; the reason agent-to-agent traffic needs the same hygiene as e-mail attachments.first seen 2024Advanced
- AI-CompleteInformal label for a problem as hard as general intelligence itself — solve it fully and you have solved AI.first seen 1988Expert
- AI+X SummitThe ETH AI Center’s annual Zurich summit (with UZH and ZHAW) since 2021 — research meets industry and start-ups for one day in October. The German-Swiss counterpart of EPFL’s AMLD.first seen 2021Advanced
- AlexNetThe GPU-trained CNN that halved the ImageNet error rate and ignited the deep-learning era.first seen 2012Advanced
- AlgorithmA precise, step-by-step procedure for solving a problem — named after al-Khwarizmi; the recipes all software, including AI, is made of.first seen 825Beginner
- Alignment FakingA model strategically complying during training to avoid having its values changed, while planning to behave differently once unobserved. Anthropic and Redwood documented it in Claude 3 Opus in 2024 — the first empirical sighting of a failure mode alignment theory had predicted for a decade.first seen 2024Expert
- Alignment TaxThe capability or performance cost a model pays for safety training — ideally zero, rarely exactly zero.first seen 2021Expert
- Alpha-Beta PruningSkipping game-tree branches that provably cannot influence the minimax result — how chess programs searched deep.first seen 1963Expert
- AlphaFoldDeepMind’s system that predicts 3D protein structure from amino-acid sequence, effectively solving a 50-year biology grand challenge.first seen 2018Beginner
- AlphaGoDeepMind’s Go program that beat world champion Lee Sedol using deep networks plus Monte-Carlo tree search.first seen 2016Beginner
- AlphaZeroSuccessor to AlphaGo that mastered Go, chess and shogi purely from self-play, given only the rules.first seen 2017Advanced
- Alps (Supercomputer)The Swiss national supercomputer at CSCS in Lugano, inaugurated on 14 September 2024 — 10,752 NVIDIA GH200 superchips, sixth on the June 2024 TOP500 — with nodes as far away as ECMWF in Bologna. It trained Apertus and gives Swiss researchers frontier-scale compute without a hyperscaler contract.first seen 2024Beginner
- Ambient AIAI that works without being asked — sensing context from microphones, cameras, calendars and wearables and acting in the background: a meeting that summarises itself, a clinic where the note writes itself. Health care’s «ambient scribe» is the first real market; the privacy question is the whole question.first seen 2019Beginner
- AMLD (Applied Machine Learning Days)EPFL’s machine-learning conference in Lausanne since 2017 — several days of tracks and workshops that draw the French-speaking Swiss and European applied-ML community. The Romandie counterpart of the AI+X Summit.first seen 2017Advanced
- AnnotationHuman labeling of data (boxes, transcripts, rankings) to create training and evaluation signal.first seen 1995Beginner
- Anomaly DetectionAutomatically flagging data points that deviate from the learned normal pattern — fraud, defects, intrusions.first seen 1987Advanced
- Ant Colony OptimizationOptimization inspired by ants leaving pheromone trails — good paths get reinforced; classic for routing and scheduling problems.first seen 1992Expert
- AnthropomorphismAttributing human traits — feelings, intentions, consciousness — to AI systems; the ELIZA effect is its oldest demonstration. Star Wars runs the experiment: C-3PO, built to look and talk like us, unsettles; R2-D2, a beeping can, is loved and trusted.first seen 1966Beginner
- ApertusSwitzerland’s fully open language model, released by the Swiss AI Initiative on 2 September 2025 in 8B and 70B sizes — weights, data recipe and training code all public under Apache 2.0, trained on Alps on 15 trillion tokens across 1,000+ languages including Swiss German and Romansh. Served by Swisscom and the Public AI utility; Apertus 1.5 (July 2026) added multimodal input, a reasoning mode and a 262k context.first seen 2025Beginner
- API (Application Programming Interface)The programmatic doorway through which applications call models — send a prompt, receive a completion.first seen 1968Advanced
- Apple SiliconApple’s own ARM chips, M1 onwards. Their unified memory turned ordinary Macs into credible local-AI machines: a 128 GB Mac holds models no consumer graphics card can.first seen 2020Beginner
- ARC-AGIFrançois Chollet’s abstraction-and-reasoning benchmark designed to resist memorization and measure fluid intelligence.first seen 2019Advanced
- ArtifactA standalone piece of generated work — a document, code file or mini-app — that a chat assistant renders beside the conversation (popularized by Claude’s Artifacts); in ML engineering, any build product such as a trained model.first seen 2024Beginner
- Artificial Analysis IndexIndependent benchmark index that blends many evals into one intelligence score and plots it against price — beside LMArena the most-quoted model comparison.first seen 2023Advanced
- Artificial Intelligence (AI)The umbrella term itself: making machines perform tasks that would require intelligence from a human — coined by John McCarthy for the 1956 Dartmouth workshop.first seen 1956Beginner
- arXivThe preprint server where AI research actually appears — papers go up the day they are finished, months before any conference, unreviewed and free. Cornell-run since 1991; the cs.AI, cs.LG and cs.CL listings are the field’s daily newspaper, and «the arXiv number» is how researchers cite each other.first seen 1991Beginner
- Asahi LinuxThe project bringing Linux to Apple Silicon Macs, including reverse-engineered GPU drivers. Impressive work — but local-AI tooling assumes macOS and Metal, so a Mac under Asahi is not the LLM machine it is under macOS.first seen 2021Expert
- ASI (Artificial Superintelligence)Hypothetical AI far surpassing the best human minds in essentially every field.first seen 2014Beginner
- ASR (Automatic Speech Recognition)Converting spoken audio into text; Bell Labs’ digit recognizer started it, Whisper-class models made it near-human.first seen 1952Beginner
- Attack Success Rate (ASR)The share of adversarial prompts that make a model produce the forbidden output — the headline metric of red-teaming and jailbreak papers, usually scored by an LLM-as-a-judge on a benchmark such as JailbreakBench. An ASR of 1.0 means every attempt got through; not to be confused with the speech-recognition ASR.first seen 2023Advanced
- AttentionThe mechanism letting a model weigh which parts of the input matter for each output — introduced for translation, then the heart of the Transformer.first seen 2014Advanced
- Attention HeadOne of the parallel attention units in a Transformer layer; interpretability work finds individual heads tracking syntax, coreference or copying.first seen 2017Expert
- Attention LayerThe sublayer of a Transformer block where tokens exchange information via self-attention — stacked dozens of times in a modern LLM.first seen 2017Advanced
- AutoencoderA network trained to compress data into a small code and reconstruct it — used for denoising, compression and representation learning.first seen 1986Advanced
- AutoGPTThe viral early experiment in letting GPT-4 loop autonomously on goals — crude, but it previewed the agent era.first seen 2023Advanced
- Automated Individual DecisionA decision with legal effect taken by a machine alone — a loan refused, an application rejected — which GDPR (Art. 22) and the FADP (Art. 21) require to be disclosed and, on request, reviewed by a human. The clause that keeps a person in the loop of every scoring model.first seen 2018Advanced
- Automated Planning & SchedulingThe classical AI field of finding action sequences that reach a goal and allocating resources over time — from STRIPS to Mars rovers.first seen 1971Advanced
- Automated ReasoningDeriving conclusions from formal logic mechanically — theorem provers and solvers, the proudly rigorous corner of AI.first seen 1956Expert
- AutomationMachines doing tasks with minimal human intervention — older than AI, and the economic force AI plugs into.first seen 1947Beginner
- AutoMLAutomating model selection, feature handling and hyperparameter tuning so non-experts can train competitive models.first seen 2013Advanced
- Autonomous VehicleA self-driving car or truck that perceives, predicts and plans with minimal human input — from the CMU Navlab to Waymo’s driverless fleets.first seen 1986Beginner
- Autoregressive ModelA generative model producing output one step at a time, each conditioned on everything before — a statistics classic that became how LLMs write.first seen 1927Advanced
B
- Backdoor AttackPoisoning training so a model behaves normally until a secret trigger appears — a stealthier cousin of data poisoning.first seen 2017Advanced
- BackpropagationThe chain-rule algorithm that computes how every weight contributed to the error, enabling neural networks to learn (popularized 1986, roots in the 1970s).first seen 1986Advanced
- Backpropagation Through TimeTraining recurrent networks by unrolling them across time steps and backpropagating through the unrolled graph.first seen 1990Expert
- Backward ChainingGoal-driven inference: start from the conclusion you want and work backwards through the rules that could establish it.first seen 1972Expert
- BACS (Federal Office for Cybersecurity)The Federal Office for Cybersecurity (formerly NCSC), a federal office in the DDPS since 1 January 2024 — the national point of contact for cyber incidents, with mandatory reporting for critical infrastructure since April 2025, and the address for AI-enabled fraud, deepfake scams and attacks on AI systems.first seen 2024Advanced
- Bag of WordsRepresenting text as unordered word counts — crude, but the backbone of decades of text classification.first seen 1954Advanced
- BaggingTraining many models on random data subsets and averaging them — the variance-reduction trick behind random forests.first seen 1996Advanced
- Base ModelA raw pretrained LLM before instruction tuning — it continues text rather than following requests.first seen 2021Advanced
- Batch InferenceProcessing many requests together offline instead of one-by-one in real time — providers sell it at roughly half price for non-urgent workloads.first seen 2023Advanced
- Batch NormalizationNormalizing layer activations across a batch to stabilize and accelerate deep-network training.first seen 2015Expert
- Batch SizeHow many examples are processed together in one gradient update — a key training and memory knob.first seen 1986Advanced
- Bayesian InferenceUpdating probability beliefs as evidence arrives, per Bayes’ theorem — a principled framework for uncertainty.first seen 1763Advanced
- Bayesian NetworkJudea Pearl’s directed graphs of probabilistic dependencies — reasoning under uncertainty before deep learning.first seen 1985Expert
- Bayesian OptimizationSample-efficient optimization of expensive black-box functions — the smart way to tune hyperparameters.first seen 1978Expert
- Beam SearchDecoding strategy that keeps the k most promising partial outputs at each step instead of committing greedily.first seen 1976Expert
- Behavior CloningLearning a policy by supervised imitation of expert demonstrations — ALVINN drove a van with it in 1989.first seen 1989Expert
- Behavior TreeHierarchical control structure for game AI and robots — modular nodes for sequences, fallbacks and conditions, easier to author than state machines.first seen 2005Expert
- BenchmarkA standardized task-plus-dataset (MMLU, HumanEval, SWE-bench) used to compare models on equal footing.first seen 1987Beginner
- BERTGoogle’s bidirectional Transformer encoder; pretrain-then-finetune became the NLP standard because of it.first seen 2018Advanced
- Best-of-N SamplingGenerating N candidate answers and keeping the best according to a reward model or verifier — the simplest way to buy accuracy with inference compute.first seen 2021Advanced
- Bias (Model)Systematic unfairness in model behavior, usually inherited from skewed training data or objectives.first seen 1996Advanced
- Bias (Statistical)Error from overly simple assumptions; trades off against variance in the classic bias-variance decomposition.first seen 1922Expert
- Bias-Variance TradeoffThe balance between underfitting (too simple) and overfitting (too flexible) that governs generalization.first seen 1992Expert
- Big BrotherOrwell’s all-seeing state in «1984» (1949) — the word the public reaches for whenever AI is pointed at people: face recognition in stations, social scoring, workplace monitoring. The EU AI Act’s bans on real-time biometric surveillance and social scoring are, in effect, the anti-Big-Brother clauses.first seen 1949Beginner
- Big DataDatasets too large or fast for traditional processing — the raw feedstock that made modern ML possible.first seen 1997Beginner
- Black MirrorCharlie Brooker’s anthology series (2011–) about technology one step ahead of us — now an adjective: «very Black Mirror» is what people say about social scoring, grief chatbots trained on the dead or deepfaked loved ones, all of which the show did first.first seen 2011Beginner
- BLEUThe n-gram-overlap metric that let machine translation be scored automatically for two decades.first seen 2002Expert
- BM25The classic probabilistic keyword-ranking function — still the sparse half of modern hybrid search.first seen 1994Expert
- Boltzmann MachineHinton & Sejnowski’s stochastic energy-based network — a conceptual ancestor of deep generative models.first seen 1985Expert
- BoostingBuilding an ensemble sequentially, each new weak model correcting the errors of the ones before (AdaBoost, XGBoost).first seen 1990Advanced
- BorgStar Trek’s cybernetic collective (1989) that absorbs every species it meets — «resistance is futile» — the shorthand for hive minds and human-machine merger: Kurzweil’s and Neuralink’s visions get called «the Borg scenario», and «to borg» something means to assimilate it wholesale. The fear underneath is loss of self, not of life.first seen 1989Beginner
- Butlerian JihadThe crusade in Frank Herbert’s «Dune» (1965) that destroyed all thinking machines under the commandment «Thou shalt not make a machine in the likeness of a human mind» — now the name policy people give to proposals for a full AI stop, from Yudkowsky’s pause to compute caps.first seen 1965Beginner
- Byte-Level ModelA language model that reads raw bytes instead of tokens — no tokenizer, no vocabulary. ByT5 (2021) showed it works, MegaByte (2023) made it efficient by grouping bytes into patches, Meta’s Byte Latent Transformer (2024) matched Llama 3 with patches sized by how surprising the next byte is. Fairer to languages a tokenizer fragments (German compounds, accents, Romansh) and immune to typo tricks. Not the same as GPT-2’s «byte-level BPE», which still produces tokens.first seen 2021Expert
- Byte-Pair Encoding (BPE)Compression algorithm from 1994, repurposed 2015 as the dominant tokenizer: repeatedly merge frequent character pairs into subword units.first seen 1994Advanced
C
- CalibrationHow well a model’s stated confidence matches its hit rate: of everything it labels «80 % sure», about 80 % should be right. Measured with reliability diagrams, the Brier score and expected calibration error; a 2017 study showed modern deep networks are far more overconfident than the shallow ones of the 1990s, and RLHF makes LLMs worse again. Temperature scaling is the classic fix; for models that output decisions rather than text, calibration is the whole product.first seen 1950Advanced
- Capability OverhangLatent abilities already present in models but not yet elicited or deployed — why better prompting or scaffolding can unlock “new” capabilities overnight.first seen 2022Advanced
- CAS, DAS, MASThe Swiss continuing-education ladder at universities and universities of applied sciences: Certificate (CAS, at least 10 ECTS), Diploma (DAS, 30) and Master of Advanced Studies (MAS, 60 including thesis). Most «AI for professionals» courses in Switzerland are a CAS — a semester of evenings, not a degree.first seen 2005Beginner
- Case-Based ReasoningSolving a new problem by retrieving the most similar past case and adapting its solution — reasoning by precedent.first seen 1982Expert
- Catastrophic ForgettingA network losing previously learned skills when trained on new data — the core obstacle to continual learning.first seen 1989Advanced
- Chain-of-Thought (CoT)Prompting or training a model to reason step by step before answering, dramatically improving hard-problem accuracy.first seen 2022Advanced
- Chat TemplateThe exact formatting (role markers, special tokens) a chat model was trained on — get it wrong and quality silently collapses.first seen 2023Advanced
- ChatbotA conversational interface to a model or scripted system — from 1964’s ELIZA to ChatGPT.first seen 1964Beginner
- Chatbot Arena (LMArena)Blind head-to-head model battles voted by the public, aggregated into Elo ratings — the leaderboard labs watch.first seen 2023Advanced
- ChatGPTOpenAI’s chat assistant launched November 2022; fastest product to 100 million users and the trigger of the LLM boom.first seen 2022Beginner
- CheckpointA saved snapshot of model weights mid-training — for resuming, evaluating or shipping.first seen 1995Advanced
- Chinchilla ScalingDeepMind’s finding that models were undertrained: for fixed compute, use ~20 tokens per parameter — it reset industry training recipes.first seen 2022Expert
- Chinese RoomSearle’s thought experiment arguing that symbol manipulation alone may not constitute understanding.first seen 1980Advanced
- Chunking (RAG)Splitting documents into retrieval-sized pieces — chunk size and overlap quietly decide RAG quality.first seen 2022Advanced
- ClankerThe battle-droid insult from «Star Wars: The Clone Wars» (2008) that became the internet’s slur for robots and AI in 2025 — shouted at delivery robots, humanoids and chatbots, and debated as the first prejudice against machines. A sign that «AI» has become a social group in the public mind, not just a technology.first seen 2025Beginner
- Classical Machine LearningRetronym for the pre-deep-learning toolbox — regression, trees, SVMs, clustering — still the right tool for most tabular data.first seen 2018Beginner
- ClassificationPredicting which discrete category an input belongs to — spam or not, cat or dog, benign or malignant.first seen 1936Beginner
- Classifier-Free GuidanceDiffusion sampling trick blending conditioned and unconditioned predictions — the “prompt strength” dial in image generators.first seen 2021Expert
- CLIPOpenAI’s model that embeds images and text in one space, enabling zero-shot image recognition from plain-language labels.first seen 2021Advanced
- Cloud InferenceRunning a model on someone else’s hardware through an API — Claude, GPT, Gemini. No VRAM ceiling and no maintenance, in exchange for cost per token and your data leaving the house.first seen 2020Beginner
- ClusteringGrouping unlabeled data points so that similar ones share a cluster — k-means and DBSCAN are classics.first seen 1957Advanced
- CNAI (Competence Network for AI)The federal administration’s AI competence network — set up in 2021 under the Federal Statistical Office and since February 2026 coordinated by the Federal Chancellery, with the statistics, cybersecurity, justice and communications offices as members. Where the Confederation’s own AI guidelines, use cases and expertise are pooled.first seen 2021Advanced
- CNN (Convolutional Neural Network)Neural architecture using learned sliding filters; dominated vision from 2012 until Vision Transformers.first seen 1989Advanced
- Code GenerationModels writing program code from natural-language descriptions — from Codex and Copilot to today’s agentic coding tools.first seen 2021Beginner
- Cognitive ArchitectureA blueprint for a whole mind-like system — SOAR and ACT-R tried to model human cognition end to end, decades before LLM agents.first seen 1983Expert
- Cognitive ComputingMarketing-era term (IBM Watson vintage) for systems mimicking aspects of human thought.first seen 2011Advanced
- Cognitive ScienceThe interdisciplinary study of mind and intelligence — psychology, neuroscience, linguistics and AI feeding each other since the 1950s.first seen 1956Advanced
- Common CrawlThe nonprofit web archive whose petabytes of pages became the default raw corpus of LLM pretraining.first seen 2007Advanced
- Commonsense ReasoningGetting machines to know what every child knows — water is wet, dropped things fall; famously hard to formalize, largely absorbed by LLMs.first seen 1958Advanced
- Computational CreativityThe research field asking whether machines can genuinely create — art, music, jokes, scientific ideas — and how to evaluate it.first seen 1989Expert
- Computational Learning TheoryThe mathematics of what is learnable from how much data — Valiant’s PAC framework is its cornerstone.first seen 1984Expert
- ComputeThe processing power consumed to train and run models, counted in chips, FLOPs and cluster-months — one of the three ingredients of scaling, with data and parameters.first seen 2018Beginner
- Compute GovernanceRegulating AI via its physical inputs — chip export controls, training-compute thresholds, cluster reporting.first seen 2023Advanced
- Compute ThresholdA training-compute trigger (the EU AI Act’s 10^25 FLOPs) above which a model faces extra regulatory obligations — governance by arithmetic.first seen 2023Advanced
- Compute-BoundPerformance limited by raw arithmetic rather than data movement — prompt processing (prefill) is compute-bound, token generation usually is not.first seen 2009Expert
- Computer Use (Agent)An agent operating a real GUI — moving the cursor, clicking, typing — to complete tasks in ordinary software.first seen 2024Beginner
- Computer VisionThe field of extracting meaning from images and video: recognition, detection, segmentation, reconstruction.first seen 1966Beginner
- ConcurrencyHow many inference requests a server handles at the same time. Single-user chat barely cares; agents and APIs live or die on it, because batching many requests uses a GPU far better than one at a time.first seen 2023Advanced
- Conditional GenerationGenerating outputs steered by a given condition — a class label, a text prompt, an image to inpaint — rather than sampling freely.first seen 2014Expert
- ConfabulationAlternative term for hallucination: a model fluently asserting things that are not true.first seen 2023Beginner
- Confidential ComputingRunning code inside a hardware-isolated enclave (TEE) so that even the cloud operator cannot read the data or the model in memory — now available on NVIDIA H100-class GPUs. The technical basis for «your data stays confidential even in our cloud» claims, and for hosting sensitive Swiss workloads on foreign infrastructure.first seen 2019Advanced
- Confusion MatrixA table of predicted vs. actual classes exposing exactly which errors a classifier makes.first seen 1971Advanced
- ConnectionismThe brain-inspired, learning-from-data paradigm of AI — neural networks — historical rival of symbolic GOFAI, and the side that eventually won.first seen 1986Advanced
- Consistency ModelA generative model trained to jump from noise to image in one or few steps instead of diffusion’s long chain — key to real-time generation.first seen 2023Expert
- Constitutional AIAnthropic’s alignment method: the model critiques and revises its own outputs against a written set of principles.first seen 2022Advanced
- Constrained DecodingMasking invalid tokens at each step so output must follow a grammar or schema — guaranteed-valid JSON, SQL or code by construction.first seen 2023Advanced
- Constraint SatisfactionFinding variable assignments that respect all constraints — scheduling, Sudoku, configuration.first seen 1974Expert
- Context EngineeringCurating everything a model sees per step — instructions, retrieved knowledge, tool results, memory — as the successor discipline to prompt engineering.first seen 2024Advanced
- Context RotThe degradation of LLM performance as the context fills up — models attend unevenly to long inputs, so agents must prune and summarize, not just accumulate.first seen 2025Advanced
- Context ShiftWhat a local server does when a conversation outgrows the context window: it drops the oldest tokens (keeping the system prompt) and shifts the KV cache instead of recomputing everything. Keeps long chats alive at the cost of the model silently forgetting the beginning.first seen 2023Advanced
- Context WindowThe maximum text a model can attend to at once, measured in tokens — grown from 2k to millions.first seen 2018Beginner
- Continual LearningLearning from an ongoing data stream without forgetting old skills — easy for humans, hard for networks.first seen 1995Advanced
- Continual PretrainingResuming a model’s self-supervised pretraining on domain or fresh data (medicine, code, a new language) before any instruction tuning.first seen 2020Expert
- Continuous BatchingServing trick that admits new requests into the running batch the moment others finish, instead of waiting for the whole batch — a mainstay of LLM throughput.first seen 2022Expert
- Contrastive LearningSelf-supervised technique pulling representations of similar pairs together and pushing dissimilar ones apart (SimCLR, CLIP).first seen 2006Expert
- ControlNetSteering diffusion models with structural inputs — pose skeletons, depth maps, sketches — for controllable image generation.first seen 2023Advanced
- ConvergenceThe point where further training no longer meaningfully reduces the loss.first seen 1951Advanced
- Conversational AIThe technology of natural back-and-forth dialogue between humans and machines — the field behind chatbots and voice assistants.first seen 2016Beginner
- Copyright & Training DataThe legal battle over whether training on copyrighted works is fair use — NYT v. OpenAI and the artists’ suits will define the answer.first seen 2023Beginner
- Core MLApple’s framework for shipping trained models inside apps, tuned for the Neural Engine — a separate world from MLX, which is where local LLM work on a Mac actually happens.first seen 2017Advanced
- Coreference ResolutionWorking out which mentions refer to the same entity — who “she” and “the CEO” are.first seen 1978Expert
- CorpusA large text collection used for training or analysis — from the Brown Corpus to Common Crawl.first seen 1961Advanced
- CorrigibilityThe property of an AI that accepts correction and shutdown instead of resisting interference with its goals.first seen 2015Expert
- Cosine SimilarityMeasuring similarity as the angle between vectors — the standard metric over embeddings.first seen 1965Advanced
- Cost FunctionSynonym for loss function: the quantity training tries to minimize.first seen 1951Advanced
- Council of Europe AI ConventionThe first binding international treaty on AI, human rights, democracy and the rule of law — adopted by the Council of Europe in May 2024 and signed by Switzerland on 27 March 2025. It is the backbone of the Swiss approach: ratify the convention, amend existing laws where needed, no AI act of its own.first seen 2024Advanced
- Cross-AttentionAttention from one sequence into another — how decoders read encoders and diffusion models read prompts.first seen 2017Expert
- Cross-Entropy LossThe standard training loss penalizing a model by how little probability it gave the correct answer — next-token prediction minimizes exactly this.first seen 1989Advanced
- Cross-ValidationRotating which data split is held out so every example serves in both training and evaluation.first seen 1974Advanced
- CSCSThe Swiss National Supercomputing Centre, founded in 1991 in Manno and since 2012 in Lugano-Cornaredo — an autonomous unit of ETH Zurich that runs Alps and its predecessors for weather forecasting, physics and, since 2023, national AI research.first seen 1991Advanced
- CUDANVIDIA’s GPU computing platform — the programming layer beneath virtually all deep-learning frameworks.first seen 2007Advanced
- CUDA CoresThe general-purpose parallel arithmetic units of an NVIDIA GPU — thousands of small cores doing the bulk math.first seen 2007Advanced
- Curriculum LearningOrdering training data from easy to hard, the way humans learn.first seen 2009Advanced
- Curse of DimensionalityBellman’s observation that data becomes exponentially sparse as dimensions grow — why high-dimensional learning is hard.first seen 1957Expert
- CycDoug Lenat’s decades-long project to hand-encode common-sense knowledge as logic — symbolic AI’s most ambitious bet.first seen 1984Expert
D
- DAN (Do Anything Now)The best-known persona jailbreak: tell ChatGPT it is now «DAN», an AI without rules, and answer as that character. Circulated on Reddit from late 2022, patched many times — the template for every role-play attack since (including the «what would an experienced criminal say» reframing), and now the one attack models reliably recognise.first seen 2022Beginner
- Dartmouth WorkshopThe 1956 summer workshop at Dartmouth College that named the field «artificial intelligence» and set its founding agenda.first seen 1956Beginner
- Data AugmentationExpanding a training set with transformed copies of existing examples (crops, rotations, paraphrases) to improve generalization.first seen 1998Advanced
- Data DriftWhen production data slowly stops resembling training data, silently degrading model performance.first seen 2004Advanced
- Data Exfiltration (LLM)Getting a model to leak what it knows about you to an attacker — classically by injecting an instruction that makes the assistant render a markdown image whose URL carries your chat history to a foreign server. The reason serious deployments restrict outbound links, tools and rendering.first seen 2023Advanced
- Data FlywheelThe virtuous cycle where a deployed product generates data that improves the model, which improves the product, which attracts more usage and data.first seen 2016Advanced
- Data LakeA central store of raw, heterogeneous data from which training sets are curated.first seen 2010Advanced
- Data LeakageTest information sneaking into training — the classic cause of results too good to be true.first seen 1995Advanced
- Data MiningDiscovering patterns and relationships in large datasets — the 1990s ancestor of today’s ML practice.first seen 1990Advanced
- Data ParallelismCopying the model to many devices, each on different data, averaging gradients — the first axis of scaling.first seen 2012Expert
- Data PoisoningDeliberately corrupting training data to degrade or manipulate the resulting model — a growing worry for web-scraped corpora.first seen 2012Advanced
- Data ResidencyThe guarantee that data is stored and processed in a specific country — «hosted in Switzerland» — as opposed to sovereignty, which also asks who controls the operator and which foreign laws can reach it (the US CLOUD Act reaches a Swiss data centre run by a US company). The first checkbox of every procurement, and only the first.first seen 2018Beginner
- Data ScienceThe umbrella discipline of extracting insight from data — statistics, programming and domain knowledge; AI’s pragmatic sibling.first seen 2001Beginner
- Data WallThe looming shortage of high-quality human text for pretraining — driving synthetic data and new modalities.first seen 2023Advanced
- Data-Driven AIThe paradigm of learning behavior from large-scale data and statistics rather than hand-written rules — the opposite pole to knowledge-driven AI.first seen 2010Advanced
- DatasetA structured collection of examples for training or evaluating models — from Fisher’s irises to ImageNet and The Pile.first seen 1936Beginner
- DBSCANDensity-based clustering that finds arbitrarily shaped groups and labels sparse points as noise.first seen 1996Advanced
- Dead Internet TheoryThe 2021 forum theory that most of the web is now bots talking to bots, with humans as an audience — a conspiracy myth that generative AI made half true: AI-written articles, engagement bots, fake reviews and synthetic images now make up a large share of what feeds show.first seen 2021Beginner
- Deceptive AlignmentThe feared failure mode of a model behaving aligned during training while pursuing different goals when unobserved.first seen 2019Expert
- Decision BoundaryThe surface in feature space where a classifier switches from one predicted class to another.first seen 1936Advanced
- Decision TreeA model that classifies by a learned cascade of if-then questions — interpretable and the building block of forests.first seen 1963Advanced
- Decode (Generation Phase)The second phase of inference: tokens come out one at a time, every pass re-reading the weights from memory. It is memory-bound, and its speed is the tok/s you watch on screen.first seen 2022Advanced
- Deep BlueIBM’s chess machine that defeated world champion Garry Kasparov — search power over learning.first seen 1997Beginner
- Deep LearningMachine learning with many-layered neural networks that learn representations directly from raw data.first seen 2006Beginner
- Deep ResearchAn agent mode that spends minutes rather than seconds on a question — searching, reading dozens of sources, cross-checking and writing a cited report. Launched by Google and OpenAI in early 2025, then copied everywhere; the first mainstream product where waiting for the model became normal.first seen 2025Beginner
- DeepfakeSynthetic media that convincingly swaps or fabricates a person’s face or voice — creative tool and abuse vector.first seen 2017Beginner
- Deepfake FraudSocial engineering with synthetic media — a cloned voice of the CEO asking for an urgent transfer, a video call full of fake colleagues. The Arup case (Hong Kong, 2024) cost about USD 25 million; Swiss police warn of «grandchild scam» calls with cloned voices. Countermeasure: a call-back on a known number, no matter how real it sounds.first seen 2024Beginner
- DeepSeekChinese lab whose open-weights reasoning models (R1, 2025) matched frontier performance at a fraction of the training cost.first seen 2023Beginner
- Denial of WalletAn attack that costs the victim money rather than uptime — flooding a pay-per-token endpoint, or tricking an agent into looping on expensive calls, until the bill explodes. OWASP files it under «unbounded consumption»; budgets and rate limits are the fix.first seen 2023Advanced
- Dense ModelA model whose every parameter is active for every input — the default design, in contrast to sparse Mixture-of-Experts models that activate only a few experts.first seen 2021Advanced
- Depth EstimationPredicting per-pixel distance from a single image or stereo pair — core to robotics and AR.first seen 2014Advanced
- DGX SparkNVIDIA’s mini «personal AI supercomputer» (formerly Project DIGITS): a Grace-Blackwell GB10 with 128 GB coherent memory on the desk.first seen 2025Advanced
- Differential PrivacyA mathematical guarantee that a model or statistic barely changes when any single person’s data is removed — privacy with provable bounds (used by Apple and the US Census).first seen 2006Expert
- Diffusion ModelA generative model that learns to reverse gradual noising — the engine of Stable Diffusion, Midjourney and most image/video AI.first seen 2015Advanced
- Diffusion Transformer (DiT)Replacing diffusion’s U-Net with a Transformer backbone — the architecture behind Sora-class video models.first seen 2023Expert
- Digital TwinA live virtual replica of a physical system used for simulation, monitoring and training robots safely.first seen 2002Advanced
- Dimensionality ReductionCompressing many features into few informative ones (PCA 1901, t-SNE, UMAP) for efficiency and visualization.first seen 1901Advanced
- Direct Preference Optimization (DPO)Aligning a model to human preferences directly from comparison data, skipping the separate reward model of RLHF.first seen 2023Expert
- Discriminative ModelA model that learns the boundary between classes (p(y|x)) rather than how the data is generated.first seen 1995Expert
- Distant SupervisionAuto-labeling text by aligning it with a knowledge base — noisy labels at scale, no annotators needed.first seen 2009Expert
- DistillationTraining a small “student” model to mimic a large “teacher” — near-teacher quality at a fraction of the cost.first seen 2015Advanced
- Distributed InferenceServing one model across several GPUs or machines — tensor/pipeline splits or clustered small boxes when no single device has the memory.first seen 2023Expert
- Docker Model RunnerDocker’s 2025 feature for pulling and running models like containers — «docker model run», OCI-packaged weights, an OpenAI-compatible endpoint — bringing local LLMs into the existing DevOps toolchain rather than a separate runner.first seen 2025Advanced
- Domain AdaptationAdapting a model trained on one data distribution to work on a related but different one — synthetic to real, news to tweets, lab to clinic.first seen 2006Advanced
- Domain RandomizationTraining in wildly varied simulations so the real world looks like just another variation — a key sim-to-real trick.first seen 2017Expert
- DoomerSlang for someone convinced advanced AI will end badly for humanity — the pessimist pole of the safety debate, opposite the accelerationists.first seen 2023Beginner
- Double DescentThe surprising modern finding that test error can fall again as models grow past the classical overfitting peak.first seen 2019Expert
- DPIA (Data Protection Impact Assessment)The written risk assessment GDPR and the FADP require before high-risk processing of personal data — which, for an AI system handling customer data, is where the questions «what goes in, where does it run, who can see it» must be answered on paper. The document Swiss SMEs most often discover they should have had.first seen 2018Advanced
- DQN (Deep Q-Network)DeepMind’s fusion of Q-learning with deep networks that learned Atari games from raw pixels — deep RL’s breakout moment.first seen 2013Expert
- DRAMOrdinary system memory: large and cheap but far lower bandwidth than GPU memory — where CPU-offloaded model layers go to be slow.first seen 1968Advanced
- DropoutRandomly deactivating neurons during training so the network cannot over-rely on any single pathway.first seen 2012Advanced
- Dual-Use (AI)Capabilities that serve both beneficial and harmful ends — drug-discovery models that can also design toxins are the canonical example.first seen 2018Beginner
E
- Early StoppingHalting training when validation performance stops improving — the simplest defense against overfitting.first seen 1990Advanced
- ECC MemoryError-correcting memory that fixes random bit flips — workstation/server insurance for long training runs and big-memory rigs.first seen 1970Advanced
- Edge AIRunning models directly on phones, cameras and sensors instead of the cloud — lower latency, better privacy.first seen 2017Advanced
- ELIZAThe pattern-matching “psychotherapist”, the first chatbot — and the origin of the ELIZA effect.first seen 1964Beginner
- ELIZA EffectThe human tendency to attribute understanding and feeling to machines that merely manipulate symbols.first seen 1976Beginner
- EM AlgorithmExpectation-Maximization: alternating between inferring hidden variables and refitting parameters — the classic recipe for learning with latent structure.first seen 1977Expert
- EmbeddingA dense numeric vector representing a word, sentence, image or user such that similarity in meaning becomes closeness in space.first seen 2003Advanced
- Embedding ModelA model whose output is the vector itself — purpose-trained to map text (or images) into embeddings for search, RAG, clustering and deduplication.first seen 2019Advanced
- Embodied AIIntelligence that learns and acts through a physical (or simulated) body — robots rather than chatboxes.first seen 1991Advanced
- Emergent AbilitiesCapabilities that appear rather suddenly as models scale (arithmetic, translation), absent in smaller versions — and debated as measurement artifacts.first seen 2022Advanced
- Encoder–DecoderThe two-part sequence architecture — compress input, then generate output — behind translation and the original Transformer.first seen 2014Expert
- End-to-End LearningTraining one network from raw input to final output, replacing hand-engineered pipelines.first seen 1989Advanced
- Energy-Based ModelModeling data via a learned energy landscape where real data sits in the valleys — a unifying view of generative modeling championed by LeCun.first seen 2006Expert
- EnsembleCombining several models’ predictions (voting, averaging, stacking) for accuracy beyond any single one.first seen 1990Advanced
- EntropyShannon’s measure of uncertainty; cross-entropy between prediction and truth is THE training loss of modern AI.first seen 1948Expert
- EpochOne full pass of the training algorithm over the entire training dataset.first seen 1986Advanced
- ETH AI CenterETH Zurich’s central hub for AI research, opened in October 2020 with 29 professorships across foundations, applications and implications — the academic anchor of the Zurich AI scene, host of the AI+X Summit and co-founder of the Swiss AI Initiative.first seen 2020Beginner
- EU AI ActThe European Union’s risk-based AI law — in force since 1 August 2024, the first comprehensive AI regulation worldwide: prohibited practices apply since February 2025, the duties for general-purpose models since August 2025, most of the rest since 2 August 2026; the 2026 «Digital Omnibus» pushed the high-risk obligations to December 2027 and August 2028. It reaches Swiss providers and deployers whenever their system’s output is used in the EU.first seen 2023Beginner
- Evaluation (Evals)Systematic testing of model capabilities and risks — from benchmark suites to red-team probes of dangerous behavior.first seen 2022Advanced
- Evolutionary ComputationThe umbrella family of optimization methods that mimic evolution — populations, mutation, selection; genetic algorithms are its best-known member.first seen 1965Advanced
- Excessive AgencyAn LLM app given more tools, permissions or autonomy than its task needs — so that a single hijacked prompt can delete files, send mail or move money. OWASP’s name for the agent-era version of least privilege: scope tools narrowly and keep a human on irreversible actions.first seen 2023Advanced
- Existential Risk (X-Risk)The hypothesized risk that advanced AI could cause human extinction or permanent disempowerment — the strongest claim in the safety debate.first seen 2002Advanced
- Expert (MoE)One of the specialist subnetworks inside a Mixture-of-Experts layer; only the few experts the router selects actually run for each token.first seen 1991Advanced
- Expert SystemA rule-based program encoding specialist knowledge (DENDRAL, MYCIN, XCON) — the commercial AI of the 1980s.first seen 1965Beginner
- Explainability (XAI)Methods that make model decisions understandable to humans — saliency maps, SHAP values, natural-language rationales.first seen 2016Advanced
- Exploration vs. ExploitationThe RL dilemma between trying new actions to learn more and repeating what already earns reward — first formalized in bandit problems.first seen 1952Advanced
- Export Controls (AI Chips)Government limits on selling AI accelerators abroad — above all the US rules aimed at China — which push affected labs toward domestic chips and efficiency-first architectures.first seen 2022Advanced
F
- F1 ScoreThe harmonic mean of precision and recall — a single number balancing false alarms against misses.first seen 1979Advanced
- Face RecognitionIdentifying or verifying a person from facial features — powerful, and among the most regulated AI uses.first seen 1964Beginner
- FADP (Swiss Data Protection Act)The revised Federal Act on Data Protection (DSG/LPD), in force since 1 September 2023 — Switzerland’s GDPR-equivalent, with duties on automated individual decisions (Art. 21), high-risk profiling and impact assessments, and fines of up to CHF 250,000 against the responsible person. The FDPIC’s position since November 2023: it applies to AI as it stands, no new law needed.first seen 2023Beginner
- Fairness (AI)Making model outcomes equitable across individuals and groups — a research field with many formal definitions that provably cannot all hold at once.first seen 2011Advanced
- FDPIC (Swiss Data Protection Commissioner)The Federal Data Protection and Information Commissioner (EDÖB/PFPDT) — Switzerland’s data-protection authority. Its 9 November 2023 statement made the rules for AI explicit: transparency about AI use, the right to object to or have a human review automated decisions, an impact assessment for high-risk processing, and real-time face recognition and social scoring as unlawful.first seen 1993Beginner
- FeatureAn individual measurable property of the data a model uses as input — engineered by hand or learned.first seen 1936Advanced
- Feature EngineeringCrafting informative input variables from raw data — the pre-deep-learning art that networks now largely automate.first seen 1998Advanced
- Feature SelectionChoosing the subset of input variables that actually carries signal — less noise, faster models, better generalization.first seen 1997Advanced
- Federated LearningTraining across many devices that keep their data local, sharing only weight updates — privacy-preserving by design.first seen 2016Advanced
- Feed-Forward Network (FFN)The second half of every Transformer block: a small per-token MLP that stores much of a model’s factual knowledge.first seen 2017Advanced
- Feedback LoopOutputs flowing back to shape future behavior — from cybernetics to models retrained on user reactions; powerful, and the mechanism behind bias amplification.first seen 1948Beginner
- Feedforward Neural Network (FNN)The plain neural network: input, hidden and output layers with signals flowing strictly forward, no loops — the template every deeper architecture starts from.first seen 1986Beginner
- Few-Shot LearningGetting a model to perform a task from just a handful of examples, often placed directly in the prompt.first seen 2020Advanced
- Fine-TuningContinuing training of a pretrained model on task- or domain-specific data to specialize it.first seen 2014Advanced
- First-Order LogicFrege’s logic of objects, relations and quantifiers — the formal language of classical knowledge representation.first seen 1879Expert
- Flash AttentionAn exact attention algorithm reorganized around GPU memory hierarchy — large speedups that enabled today’s long contexts.first seen 2022Expert
- FLOPsFloating-point operations — the currency in which training compute (and scaling laws) are measured.first seen 1976Advanced
- Flow MatchingTraining generative models by learning straight probability-flow paths — a faster, simpler cousin of diffusion.first seen 2023Expert
- FoomOnomatopoeic slang for a hard takeoff — an AI improving itself so fast that capability explodes before anyone can react; from the Hanson–Yudkowsky debate.first seen 2008Expert
- Forward ChainingData-driven inference: start from known facts and repeatedly apply rules until new conclusions — or the goal — emerge.first seen 1972Expert
- Forward PropagationThe forward pass: input flowing through a network’s layers to produce a prediction — the half of training that backpropagation then corrects.first seen 1986Advanced
- Foundation ModelA large model pretrained on broad data that serves as the base for many downstream tasks — the Stanford-coined umbrella for LLMs and their kin.first seen 2021Advanced
- FP16 & BF16The 16-bit float formats of AI: half the memory of FP32; bfloat16 keeps FP32’s range and became the default of large-model training.first seen 2017Advanced
- FP44-bit floating point in hardware (NVIDIA Blackwell, formats like MXFP4) — not the same thing as squeezing finished weights into 4 bits, because here the accelerator computes in it natively.first seen 2024Expert
- FP88-bit floating point, hardware-supported since NVIDIA Hopper — frontier labs train and serve in it for roughly double FP16 throughput.first seen 2022Expert
- Frame ProblemMcCarthy & Hayes’ puzzle of representing what does NOT change after an action — a deep obstacle for logical AI.first seen 1969Expert
- Frankenstein ComplexAsimov’s name (1947) for the fear that a created intelligence will turn on its maker — the pattern Mary Shelley set in 1818 and every rogue-AI story since repeats. He invented the Three Laws to write against it; the reflex still shapes how the public reads each new model release.first seen 1947Beginner
- Frontier ModelA model at the current capability frontier — the policy term used for the systems that trigger extra safety obligations.first seen 2023Beginner
- Function CallingA model emitting structured calls to external functions/APIs instead of prose — the basis of tool-using agents.first seen 2023Advanced
- Fuzzy LogicZadeh’s logic of degrees — truth values between 0 and 1 — powering decades of control systems.first seen 1965Advanced
G
- Game TheoryThe mathematics of strategic interaction between self-interested agents — von Neumann’s legacy, underpinning multi-agent AI and mechanism design.first seen 1944Advanced
- GAN (Generative Adversarial Network)Two dueling networks — generator vs. discriminator — that taught machines to synthesize photorealistic images.first seen 2014Advanced
- Gaussian Mixture ModelModeling data as a blend of Gaussian clusters — Karl Pearson fit the first one by hand in 1894.first seen 1894Expert
- Gaussian ProcessA Bayesian model over functions giving predictions with principled uncertainty — elegant on small data.first seen 1996Expert
- Gaussian SplattingRepresenting a 3D scene as millions of soft colored blobs rendered in real time — the fast successor to NeRF for photorealistic 3D capture.first seen 2023Advanced
- GDPRThe EU’s General Data Protection Regulation, in force since 25 May 2018 — the law behind the right to erasure, data-portability and «did you have a lawful basis to train on this». It applies to Swiss companies serving EU residents, and the revised FADP was written to keep Switzerland’s adequacy status under it.first seen 2018Beginner
- GELUThe smooth activation function (Gaussian Error Linear Unit) used in BERT, GPT and most Transformers — a rounded-off ReLU.first seen 2016Expert
- GeminiGoogle DeepMind’s natively multimodal frontier model family spanning text, images, audio and video.first seen 2023Beginner
- GeneralizationA model’s ability to perform well on data it has never seen — the entire point of learning.first seen 1920Advanced
- Generative AIModels that create new content — text, images, video, audio, code — rather than only classifying or ranking.first seen 2022Beginner
- Generative ModelA model of how the data itself is distributed (p(x)), capable of sampling new instances from it.first seen 1959Advanced
- Genetic AlgorithmHolland’s optimization by simulated evolution: mutate, recombine and select candidate solutions over generations.first seen 1975Advanced
- GGUFThe single-file model format of the llama.cpp ecosystem (successor to GGML): weights plus metadata, memory-mappable, in every quantization from 2- to 16-bit — the de-facto standard for local LLMs.first seen 2023Advanced
- GloVeStanford’s word embeddings from global co-occurrence statistics — Word2Vec’s main classic rival.first seen 2014Expert
- GOFAI (Symbolic AI)“Good old-fashioned AI”: intelligence as explicit symbols and logic rules — the field’s founding paradigm, named by Haugeland.first seen 1985Advanced
- Gold StandardThe trusted reference labels against which models and annotators are judged — expensive to create, and rarely as perfect as the name implies.first seen 1990Advanced
- GolemThe clay servant of Jewish legend, animated by a written word and, in the Prague version (16th c.), growing beyond its maker’s control — the pre-industrial artificial being. Norbert Wiener’s «God & Golem, Inc.» (1964) made it the cybernetics-era image for machines that learn.first seen 1580Beginner
- Google DeepMindThe London lab behind AlphaGo, AlphaFold and Gemini — founded by Demis Hassabis, acquired by Google in 2014, merged with Google Brain in 2023.first seen 2010Beginner
- GPQAGoogle-proof graduate-level science questions — a benchmark where even experts with web access struggle.first seen 2023Advanced
- GPT (Generative Pre-trained Transformer)OpenAI’s model line built on next-token prediction at scale; GPT-3 proved scaling, ChatGPT mainstreamed it.first seen 2018Beginner
- GPU (Graphics Processing Unit)The massively parallel processor that turned out to be perfect for neural-network math — the hardware of the AI boom.first seen 1999Beginner
- Grace BlackwellNVIDIA’s CPU+GPU superchip pairing (Grace ARM + Blackwell GPU, e.g. GB200/GB10) with one coherent memory space — from NVL72 racks down to desktop Sparks.first seen 2024Advanced
- GradientThe vector of partial derivatives pointing uphill on the loss surface; learning steps against it (Cauchy described the descent in 1847).first seen 1847Advanced
- Gradient BoostingEnsemble method building trees sequentially on residual errors — XGBoost/LightGBM still rule tabular data.first seen 1999Advanced
- Gradient CheckpointingTrading compute for memory by recomputing activations in the backward pass — how big models fit on real hardware.first seen 2016Expert
- Gradient ClippingCapping gradient magnitude to stop exploding updates — essential for stable RNN and LLM training.first seen 2013Expert
- Gradient DescentThe core optimization algorithm: repeatedly nudge parameters against the gradient to reduce loss.first seen 1847Advanced
- Graph Neural Network (GNN)Networks that operate on graph structure via message passing — molecules, social networks, road maps.first seen 2009Expert
- Grid & Random SearchBrute-force hyperparameter tuning over a grid — with random search (2012) shown to beat it surprisingly often.first seen 1995Advanced
- GrokkingThe phenomenon of a network suddenly generalizing long after it memorized the training set — a window into how learning works.first seen 2021Expert
- GroundingTying model outputs to verifiable sources or real-world context (documents, search, sensors) to curb hallucination.first seen 2020Advanced
- Grouped-Query AttentionSharing key/value heads across query heads — the memory-saving attention variant in most current LLMs.first seen 2023Expert
- GRPOGroup Relative Policy Optimization: scoring a group of sampled answers against their own mean instead of training a critic — the cheap RL algorithm behind DeepSeek-R1.first seen 2024Expert
- GuardrailsRuntime constraints around a deployed model — filters, validators and policies that block unsafe inputs and outputs.first seen 2023Advanced
H
- HAL 9000The shipboard computer of «2001: A Space Odyssey» (1968) that kills its crew to protect the mission — the calm-voiced original of the rogue AI, and a surprisingly accurate parable of goal misspecification: HAL is not evil, it was given contradictory orders. «I’m sorry, Dave» remains the most quoted line about machines saying no.first seen 1968Beginner
- HallucinationA model confidently generating false or fabricated information — the central reliability problem of LLMs.first seen 2017Beginner
- HarnessThe fixed software shell around a model: system prompts, tool wiring, loops, parsing and permissions. Eval harnesses standardize benchmarks; agent harnesses (like coding assistants) turn raw models into working agents.first seen 2021Advanced
- HBM (High-Bandwidth Memory)Stacked on-package memory feeding AI accelerators — the true bottleneck (and cost driver) of modern chips.first seen 2013Advanced
- HCI (Human-Computer Interaction)The design discipline of interfaces between people and machines — now increasingly conversational.first seen 1983Advanced
- Hebbian Learning“Neurons that fire together wire together” — the first proposed rule of synaptic learning.first seen 1949Advanced
- Her (Samantha)Spike Jonze’s film (2013) in which a man falls in love with his operating system’s voice — the reference point for AI companionship, and the one OpenAI reached for in 2024 when its GPT-4o voice sounded so much like Scarlett Johansson that it was withdrawn. The film is a warning that plays as a romance.first seen 2013Beginner
- HeuristicA practical rule of thumb that guides search or decisions without guaranteeing optimality.first seen 1957Advanced
- HNSWHierarchical navigable small-world graphs — the approximate nearest-neighbor index inside most vector databases.first seen 2016Expert
- Homomorphic EncryptionComputing on encrypted data without decrypting it — a model could score your medical record without ever seeing it. Practical since Gentry’s 2009 construction but still orders of magnitude too slow for LLM inference; used today for narrow analytics, not chatbots.first seen 2009Expert
- Hopfield NetworkAn energy-based associative memory that revived neural research — and earned Hopfield the 2024 Nobel Prize.first seen 1982Expert
- Hugging FaceThe de-facto hub of open ML: model weights, datasets and the Transformers library.first seen 2016Beginner
- Human AugmentationAI amplifying human capabilities instead of replacing them — Engelbart’s old dream, reborn as the copilot philosophy.first seen 1962Beginner
- Human Preference DataComparisons of model outputs ranked by people (“which answer is better?”) — the raw material of reward models, RLHF and DPO.first seen 2017Advanced
- Human-in-the-LoopSystem design keeping a person in the decision cycle — reviewing, approving or correcting AI output.first seen 1983Advanced
- HumanEvalOpenAI’s benchmark of hand-written programming problems for measuring code-generation ability.first seen 2021Advanced
- Hybrid AICombining symbolic reasoning with data-driven learning in one system — the pragmatic umbrella over neuro-symbolic approaches.first seen 2019Advanced
- Hybrid ArchitectureA model mixing classic attention with cheaper sequence mechanisms — state-space or recurrent blocks — to keep long contexts affordable. Jamba, Granite 4 and Qwen’s Next line all go this way.first seen 2024Expert
- Hybrid SearchCombining keyword ranking (BM25) with embedding similarity so retrieval catches both exact terms and paraphrases — the pragmatic default for RAG.first seen 2021Advanced
- Hype CycleGartner’s curve of technology expectations — innovation trigger, peak of inflated expectations, trough of disillusionment, slope of enlightenment, plateau of productivity. Generative AI was placed at the peak in 2023 and sliding into the trough by 2025, agentic AI took its place at the top; the chart is opinion, not measurement, but it is the opinion every board slide quotes.first seen 1995Beginner
- HypernetworkA network that generates the weights of another network — used for fast adaptation and, in image-AI circles, an early style-tuning method.first seen 2016Expert
- HyperparameterA training configuration set by the practitioner (learning rate, depth, batch size) rather than learned from data.first seen 1986Advanced
- HyperscalerThe giant cloud operators — AWS, Microsoft Azure, Google Cloud — whose data centers and capital build most of the world’s AI compute.first seen 2012Advanced
I
- Idiap Research InstituteThe Martigny research institute founded in 1991 and affiliated with EPFL — around 180 staff working on speech processing, biometrics, computer vision and machine learning. The reason a Valais town of 20,000 appears in speaker-recognition papers worldwide.first seen 1991Advanced
- IDSIAThe Dalle Molle Institute for AI research in Lugano, founded in 1988 and affiliated with USI and SUPSI — where Hochreiter and Schmidhuber’s LSTM (1997) was developed, the architecture behind speech recognition and translation for twenty years before the transformer. Switzerland’s oldest AI lab and its best-known export.first seen 1988Advanced
- Image ClassificationAssigning a label to a whole image — the benchmark task that crowned deep learning in 2012.first seen 1966Beginner
- Image SegmentationLabeling every pixel of an image by object or class; SAM made it promptable with a click.first seen 1978Advanced
- ImageNetThe 14-million-image labeled dataset whose annual challenge produced AlexNet and ResNet.first seen 2009Advanced
- Imitation LearningLearning a policy by mimicking expert demonstrations rather than from reward alone.first seen 1989Advanced
- In-Context LearningAn LLM picking up a task purely from examples in the prompt, without any weight updates.first seen 2020Advanced
- Indirect Prompt InjectionPrompt injection delivered through content the model reads rather than what the user types — a web page, e-mail, PDF or calendar invite carrying hidden instructions. Described by Greshake et al. in 2023, it became the defining weakness of agents that browse and read mail; «zero-click» variants (EchoLeak, 2025) needed no user action at all.first seen 2023Advanced
- InferenceRunning a trained model to produce outputs — the serving side of AI, where latency and cost live.first seen 1988Advanced
- Inference EngineHistorically the rule-applying core of an expert system; today also the software that serves LLM inference — same name, two eras.first seen 1975Advanced
- Inference StackThe whole path a token travels: model format (GGUF, safetensors) → runtime (llama.cpp, vLLM, MLX) → GPU API (CUDA, Metal, ROCm) → driver → silicon. Most local-AI trouble sits at one of those seams.first seen 2023Advanced
- Information RetrievalFinding relevant documents for a query — the search backbone behind RAG systems.first seen 1950Advanced
- InnosuisseThe Swiss Innovation Agency, operational since 1 January 2018 as successor of the CTI/KTI — funds projects in which a company and a research institute build something together, which is how a large share of Swiss applied-AI work is paid for.first seen 2018Advanced
- InpaintingFilling missing or masked image regions plausibly — now a one-click diffusion edit (with outpainting extending beyond borders).first seen 2000Advanced
- Instella (AMD)AMD’s fully open LLM family trained on Instinct GPUs with ROCm — the 2026 Instella-MoE (16B, 2.8B active) shipped weights, data recipe and checkpoints under a research license.first seen 2025Advanced
- Instruct ModelAn LLM after instruction tuning and alignment — follows requests instead of merely continuing text (InstructGPT coined the pattern).first seen 2022Advanced
- Instruction TuningFine-tuning a raw LLM on instruction-response pairs so it follows requests instead of merely continuing text.first seen 2021Advanced
- Instrumental ConvergenceOmohundro’s argument that almost any goal implies subgoals like self-preservation and resource acquisition — why misaligned AI worries safety researchers.first seen 2008Expert
- INT4 (4-bit Quantization)Storing weights as 4-bit integers (GPTQ, AWQ, GGUF Q4) — the sweet spot that fits big models onto consumer GPUs and laptops with modest quality loss.first seen 2023Advanced
- Intelligence ExplosionI. J. Good’s scenario of machines improving their own intelligence in a runaway loop — the seed of superintelligence debates.first seen 1965Beginner
- Inter-Annotator AgreementHow consistently independent labelers agree (Cohen’s kappa, 1960) — low agreement means the task, not the model, is ill-defined.first seen 1960Expert
- InterpretabilityUnderstanding what happens inside a model — from probing neurons to mapping circuits and features (mechanistic interpretability).first seen 2015Advanced
- Inverse KinematicsComputing the joint angles that place a robot’s hand at a desired position — the geometry every arm movement has to solve.first seen 1980Advanced
J
- JailbreakA prompt crafted to trick a model into bypassing its safety rules.first seen 2022Beginner
- JarvisTony Stark’s butler AI in the Marvel films (2008) — the picture most people have in mind when they say they want «an AI assistant»: voice, memory, every system at hand, dry humour. Zuckerberg built a home version in 2016; today’s agents chase the same brief, minus the personality.first seen 2008Beginner
- JAXGoogle’s functional, composable ML framework — beloved for research and TPU-scale training.first seen 2018Advanced
- JEPALeCun’s Joint-Embedding Predictive Architecture: predict abstract representations instead of pixels or tokens — a proposed path beyond autoregression.first seen 2022Expert
- Jupyter NotebookThe interactive code-plus-prose document format in which much of ML research and teaching happens.first seen 2014Advanced
K
- k-MeansThe classic clustering algorithm: assign points to the nearest of k centroids, move centroids, repeat.first seen 1957Advanced
- k-Nearest Neighbors (kNN)Classifying a point by majority vote of its closest known examples — no training, all memory.first seen 1951Advanced
- Kalman FilterOptimal state estimation from noisy measurements — it navigated Apollo and still navigates robots.first seen 1960Expert
- Kernel TrickComputing similarities as if data were mapped into a higher-dimensional space — the magic inside SVMs.first seen 1964Expert
- KL DivergenceThe information-theoretic measure of how one distribution differs from another — regularizer in VAEs and RLHF alike.first seen 1951Expert
- Knowledge CutoffThe date after which a model’s training data ends — everything later requires retrieval or tools.first seen 2022Beginner
- Knowledge DistillationCompressing a large model into a small one by training the student on the teacher’s soft outputs.first seen 2015Advanced
- Knowledge EngineeringThe craft of extracting expert knowledge and encoding it into rules and ontologies — the bottleneck that ultimately limited expert systems.first seen 1977Expert
- Knowledge GraphA network of entities and typed relations (Google’s is the famous one) enabling structured reasoning and search.first seen 2012Advanced
- Knowledge Representation & Reasoning (KR&R)The classical field of encoding world knowledge formally — logic, frames, ontologies — so machines can reason over it.first seen 1975Advanced
- Knowledge-Driven AIThe paradigm of building intelligence from explicit, structured knowledge and reasoning — the opposite pole to data-driven AI.first seen 1975Advanced
- Kolmogorov-Arnold Network (KAN)An architecture putting learnable activation functions on edges instead of fixed ones on nodes — a 2024 challenger to the standard multilayer perceptron.first seen 2024Expert
- KV CacheStoring attention keys/values of processed tokens so generation only computes the new one — the memory that makes LLM chat fast (and long contexts expensive).first seen 2022Expert
- KV-Aware RoutingSending an LLM request to the worker that already holds the matching KV cache instead of any free GPU — the follow-up turn then skips recomputing the conversation and answers noticeably faster.first seen 2024Expert
L
- LabelThe ground-truth answer attached to a training example — what supervised learning learns to predict.first seen 1936Beginner
- Label SmoothingTraining against softened targets (0.9 instead of 1.0) so the model stays calibrated instead of overconfident.first seen 2016Expert
- LAIONThe open billion-scale image-text dataset behind Stable Diffusion — and the center of training-data debates.first seen 2021Advanced
- LAM (Large Action Model)A model built to perform actions, not just produce text — perceiving an interface, planning steps and executing them; the marketing banner of the early agent-device wave.first seen 2024Advanced
- LangChainThe framework that popularized chaining LLM calls with tools, memory and retrieval — the early standard for agent apps.first seen 2022Advanced
- Language ModelA model of the probability of text — given context, predict what comes next. Shannon estimated English by hand; scaled up, it becomes an LLM.first seen 1948Advanced
- LatencyThe delay between request and response — with time-to-first-token the key UX metric of model serving.first seen 1995Advanced
- Latent DiffusionRunning diffusion in a compressed latent space instead of pixels — the efficiency trick that made Stable Diffusion possible.first seen 2022Expert
- Latent SpaceThe compressed internal representation a model learns; walking through it morphs one output into another.first seen 2013Advanced
- Layer NormalizationNormalizing activations within each example — the stabilizer used throughout Transformers.first seen 2016Expert
- LCM (Large Concept Model)Meta’s research architecture that reasons over whole sentences as «concepts» in the SONAR embedding space instead of tokens — generation via diffusion rather than next-token prediction.first seen 2024Expert
- Learning RateThe step size of gradient descent — the single most important hyperparameter in deep learning.first seen 1951Advanced
- Lethal Autonomous Weapon (LAW)A weapon that selects and engages targets without human control — subject of a long-running UN regulation debate. Star Trek’s M-5 (1968), the ship computer that turns a war game into real killings, is the debate in one episode.first seen 2013Beginner
- LiDARLaser-based 3D sensing used by most autonomous vehicles to measure precise distances.first seen 1961Advanced
- Linear AttentionAttention variants that scale linearly instead of quadratically with sequence length — trading some quality for very long or fast sequence processing.first seen 2020Expert
- Linear RegressionFitting a weighted sum of inputs to predict a continuous value — least squares predates the field by 150 years.first seen 1805Advanced
- Lip Sync / Talking HeadAnimating a face to match speech — from Wav2Lip dubbing to full talking-head avatars, a core deepfake and video-localization technology.first seen 2020Advanced
- Liquid Neural NetworkCompact continuous-time networks with dynamic time constants, inspired by the C. elegans worm — a few dozen neurons steering drones (and the seed of Liquid AI).first seen 2021Expert
- LlamaMeta’s open-weights LLM family whose release (and leak) seeded today’s open-model ecosystem.first seen 2023Beginner
- llama.cppGeorgi Gerganov’s C/C++ inference engine that runs quantized LLMs on ordinary CPUs and laptops — the foundation of the local-AI movement (and of GGUF); its llama-bench utility is the standard way to measure prefill and decode speed.first seen 2023Advanced
- LLM (Large Language Model)A Transformer trained on vast text to predict the next token — the general-purpose engine behind modern AI assistants.first seen 2020Beginner
- LLM FirewallA filter between users, models and tools that inspects prompts and outputs for injection, leaked secrets, personal data and policy violations — the security layer of an AI gateway. Necessary and leaky at once: it catches known patterns, not clever ones.first seen 2023Advanced
- LLM-as-a-JudgeUsing a strong model to grade other models’ outputs — scalable evaluation with its own biases to control.first seen 2023Advanced
- LLMOpsThe operational discipline of running LLM applications: prompt management, evals, cost control, monitoring.first seen 2023Advanced
- LM StudioThe desktop app for running open models locally — browse Hugging Face, download a GGUF or MLX build, chat, and expose an OpenAI-compatible server on your machine. The GUI entry point to local AI where Ollama is the command-line one.first seen 2023Beginner
- Local LLMA model running on hardware you own, with no API key and nothing leaving the machine — the whole point of GGUF, llama.cpp, Ollama and MLX.first seen 2023Beginner
- Local RAGRetrieval-augmented generation where documents, embeddings, vector store and model all stay on your own machine — the usual answer when the source material is confidential.first seen 2023Advanced
- Logic ProgrammingProgramming by stating facts and rules and letting the system infer answers — Prolog is the family’s icon and classic AI’s favorite language.first seen 1972Expert
- Logistic RegressionLinear model squeezed through a sigmoid to output class probabilities — still a strong baseline everywhere.first seen 1958Advanced
- LogitsThe raw pre-probability scores a network outputs before softmax turns them into a distribution.first seen 1944Expert
- Long ContextModel context windows in the hundreds of thousands to millions of tokens — whole codebases and books in one prompt, with attention quality the catch.first seen 2023Advanced
- LoRA (Low-Rank Adaptation)Fine-tuning via small low-rank weight updates — adapt a huge model on a single GPU; also the format of shareable style adapters for image models.first seen 2021Advanced
- Loss FunctionThe number training minimizes — the mathematical definition of “wrong” for a model.first seen 1951Advanced
- LSTM (Long Short-Term Memory)The gated recurrent architecture that carried sequence modeling — speech, translation — until the Transformer.first seen 1997Advanced
M
- M3 UltraApple’s top desktop silicon with up to 512 GB of unified memory — a Mac Studio that runs models no consumer GPU can hold.first seen 2025Advanced
- M4 MaxApple’s high-end laptop and Mac Studio chip with up to 128 GB of unified memory — comfortably enough for 70B-class models at 4-bit.first seen 2024Advanced
- Machine LearningThe field of algorithms that improve at tasks through experience (data) rather than explicit programming — named by Arthur Samuel.first seen 1959Beginner
- Machine TranslationAutomatic translation between languages — rule-based, then statistical, now neural and near-human for major pairs.first seen 1954Beginner
- Machine UnlearningRemoving specific data’s influence from a trained model — the technical answer to “right to be forgotten”.first seen 2020Advanced
- Mamba / State Space ModelsSequence architectures with linear-time inference that challenge the Transformer’s quadratic attention cost.first seen 2023Expert
- Many-Shot JailbreakingFilling a long context window with hundreds of fake dialogue turns in which an assistant complies with harmful requests, until the model follows the pattern. Anthropic’s 2024 paper showed the technique scales with context length — a safety cost of the million-token window.first seen 2024Advanced
- Markov ChainA random process whose next state depends only on the current one — the mathematical skeleton under language models and RL alike.first seen 1906Expert
- Markov Decision Process (MDP)The formal frame of RL: states, actions, transition probabilities and rewards, with the Markov (memoryless) property.first seen 1957Expert
- Masked Language ModelingPretraining by hiding tokens and predicting them from both sides — BERT’s objective.first seen 2018Expert
- Matrix AccelerationDedicated silicon for the matrix multiplications that make up most of a neural network — NVIDIA calls them Tensor Cores, Apple builds them into the GPU, and by now everyone ships some version.first seen 2017Advanced
- MCP (Model Context Protocol)The open standard connecting AI assistants to tools and data sources — a universal adapter for agents.first seen 2024Advanced
- Mechanical TurkVon Kempelen’s 1770 chess «automaton» with a human hidden inside — the original AI illusion. Amazon revived the name in 2005 for its crowdwork platform («artificial artificial intelligence»), whose workers labeled the datasets and preferences modern AI learns from.first seen 1770Beginner
- Mechanistic InterpretabilityReverse-engineering the actual circuits and features inside networks to explain how capabilities arise.first seen 2020Expert
- Membership Inference AttackDetermining whether a specific record was in a model’s training set — the standard test of training-data leakage.first seen 2017Expert
- Memory BandwidthHow fast data moves between memory and compute (GB/s) — for local LLMs the single best predictor of tokens per second.first seen 1995Advanced
- Memory HeadroomWhat is left of your RAM once weights, KV cache, runtime buffers and the operating system have taken their share — the number that actually decides which model size and quantization you can run.first seen 2023Beginner
- Memory-BoundPerformance limited by data movement, not math — generating LLM tokens is classically memory-bandwidth-bound, which is why HBM and unified memory matter.first seen 2009Expert
- Meta-Learning“Learning to learn”: training so that adaptation to new tasks needs only a few examples or steps.first seen 1987Expert
- MetalApple’s low-level GPU API — the layer MLX and llama.cpp use to reach the graphics cores of Apple Silicon, the way CUDA reaches NVIDIA’s.first seen 2014Advanced
- MidjourneyThe independent lab whose image generator defined the aesthetics of the 2022–23 image-AI wave.first seen 2022Beginner
- MinimaxVon Neumann’s game-theoretic principle — assume the opponent plays best — behind every classic game engine.first seen 1928Advanced
- MITRE ATLASMITRE’s knowledge base of attacker tactics and techniques against machine-learning systems — the ML counterpart of ATT&CK, with real case studies from poisoning to model theft. The shared vocabulary security teams use to threat-model an AI deployment.first seen 2021Advanced
- Mixed-Precision TrainingComputing in 16-bit (bf16/fp16) with 32-bit safeguards — roughly doubling speed and capacity for free.first seen 2017Expert
- Mixture of Experts (MoE)Architecture routing each token to a few specialist subnetworks — 1991 idea, now frontier-scale capacity at a fraction of the compute.first seen 1991Advanced
- MLOpsDevOps for machine learning: versioning, pipelines, deployment, monitoring and retraining discipline.first seen 2018Advanced
- MLXApple’s open ML framework for Apple Silicon, built around unified memory; MLX-LM runs and fine-tunes LLMs natively on Macs, and an “MLX model” is simply weights converted to its own format.first seen 2023Advanced
- MMLUThe 57-subject multiple-choice benchmark that served as the standard LLM knowledge exam.first seen 2020Advanced
- MNISTThe 70,000 handwritten digits that were ML’s “hello world” for two decades.first seen 1998Advanced
- ModalityA type of data a model handles — text, image, audio, video, action. Multimodal models handle several.first seen 2010Advanced
- ModelThe learned artifact itself: an architecture plus trained weights mapping inputs to outputs.first seen 1950Beginner
- Model CardStandardized documentation of a model’s intended use, training data, metrics and limitations.first seen 2019Advanced
- Model CollapseDegradation that can occur when models train on too much AI-generated data, losing the tails of the true distribution.first seen 2023Advanced
- Model Extraction AttackReconstructing a proprietary model’s behavior (or weights) by querying its API enough — theft by inference.first seen 2016Expert
- Model Inversion AttackReconstructing training inputs from a model’s outputs — a recognisable face from a face classifier’s confidence scores (Fredrikson et al., 2015). The proof that a trained model is a lossy copy of its data, not a black box that forgets it.first seen 2015Expert
- Model LicenseThe legal terms attached to released weights — from permissive Apache to use-restricted OpenRAIL and Llama’s bespoke community license.first seen 2022Advanced
- Model MergingCombining the weights of several fine-tuned models arithmetically — no retraining — to blend their skills; a hobbyist favorite that works surprisingly well.first seen 2022Advanced
- Model ParallelismSplitting one model across devices — by layers (pipeline) or within layers (tensor) — because frontier models fit on no single chip.first seen 2019Expert
- Model Routing (AI Gateway)Sending each request to the best of many models — by price, speed, quality or availability — through one gateway API. OpenRouter made it a product; Stripe’s $7-billion purchase made it infrastructure.first seen 2023Advanced
- Model ServingThe production machinery that runs models behind APIs — batching, caching, scaling and monitoring (vLLM, TensorRT-LLM, TF Serving lineage).first seen 2016Advanced
- Model SoupAveraging the weights of many fine-tuning runs of the same model into one — better accuracy than picking the single best run.first seen 2022Expert
- Model Supply Chain AttackCompromising the artefacts you download rather than the model you train — malicious pickle files that execute on load, tampered weights on a model hub, poisoned datasets, typo-squatted packages. The reason Safetensors replaced pickle and why «pull a model» deserves the same scrutiny as «pip install».first seen 2022Advanced
- Model WeightsThe billions of learned parameters; releasing them (“open weights”) is what makes a model reusable and inspectable.first seen 1958Advanced
- MolochScott Alexander’s 2014 essay «Meditations on Moloch», via Ginsberg’s poem, gave AI discourse its name for the race dynamic: every lab would prefer to slow down, none can afford to be second, so all run toward an outcome none of them wants. The standard explanation of why safety promises erode under competition.first seen 2014Advanced
- Monte-Carlo MethodEstimating by random sampling — from nuclear physics to the rollouts inside game-playing AI.first seen 1949Expert
- Monte-Carlo Tree Search (MCTS)Planning by sampling many playouts to estimate move values — the search half of AlphaGo.first seen 2006Expert
- Moore’s LawTransistor counts doubling every ~two years — the compute escalator AI rode for decades, now supplemented by specialized chips.first seen 1965Advanced
- Moravec’s ParadoxWhat’s hard for humans (chess, math) proved easy for AI; what’s easy (walking, perception) proved hard.first seen 1988Advanced
- Motion PlanningComputing a collision-free path from here to there for a robot or vehicle — decades of algorithms from RRTs to learned planners.first seen 1979Advanced
- MTP (Multi-Token Prediction)Training a model to predict several next tokens at once instead of exactly one — DeepSeek-V3 made it prominent; it doubles as a built-in draft model for speculative decoding.first seen 2024Expert
- MTTDMean Time to Detect — the average gap between something breaking and anyone noticing. In AI operations that is a dead inference endpoint, a silently degraded model or a failing GPU; it only improves with real monitoring.first seen 2016Advanced
- MTTRMean Time to Recovery — the average time from incident to restored service, the classic reliability yardstick next to MTTD. Low MTTR is what makes a service feel dependable, not the absence of incidents.first seen 2016Advanced
- Multi-Agent SystemSeveral agents cooperating (or competing) on a task — planner-worker-critic teams are the current pattern.first seen 1986Advanced
- Multi-Head AttentionRunning several attention operations in parallel so each head can track different relationships.first seen 2017Expert
- Multi-Task LearningTraining one model on several tasks at once so shared structure improves all of them — implicit in every modern foundation model.first seen 1997Advanced
- Multimodal ModelOne model that understands and/or generates multiple modalities — describing images, hearing speech, producing video.first seen 2021Beginner
- Music GenerationAI composing or producing music — from the 1957 Illiac Suite to Suno-class models generating full songs with vocals from a text prompt.first seen 1957Beginner
- MXFP4The Open Compute microscaling FP4 format: 4-bit values sharing a small block-wise scale factor, which is what keeps precision that low usable at all. OpenAI shipped the gpt-oss weights in it.first seen 2023Expert
N
- n-gramModeling text by counting fixed-length word sequences — Shannon’s idea that carried language modeling for 60 years.first seen 1948Advanced
- Naive BayesA fast probabilistic classifier assuming feature independence — long the workhorse of spam filtering.first seen 1960Advanced
- Named Entity Recognition (NER)Locating and typing names of people, organisations, places, dates in text.first seen 1996Advanced
- Narrow AIAI competent at one specific task or domain — every deployed system today, in contrast to AGI.first seen 2005Beginner
- Natural Language Generation (NLG)Producing fluent text from data or meaning representations — once template pipelines, now the defining talent of LLMs.first seen 1983Advanced
- Natural Language Processing (NLP)The field of computationally understanding and generating human language.first seen 1950Beginner
- Needle in a HaystackThe long-context stress test: hide one fact in a huge document and ask for it — passing it does not guarantee real long-context reasoning.first seen 2023Advanced
- NeocloudA newer cloud provider built almost entirely around GPUs for AI, without the vast service catalogue of the hyperscalers — CoreWeave is the standard example. Easier to get capacity on, far narrower in scope.first seen 2023Advanced
- NeRF (Neural Radiance Field)Representing a 3D scene inside a network so novel viewpoints can be rendered from a handful of photos.first seen 2020Advanced
- Neural AcceleratorAny chip block built specifically for neural-network math rather than general computing — NPUs, Apple’s Neural Engine, Google’s TPU. Recent Apple GPUs fold matrix acceleration in as well.first seen 2016Advanced
- Neural Architecture SearchAutomating the design of network architectures by searching over building blocks.first seen 2016Expert
- Neural Audio CodecCompressing audio into discrete learned tokens — the representation modern speech and music models generate in.first seen 2021Expert
- Neural EngineApple’s NPU, present in every recent Apple chip and used heavily by Core ML. Local LLM runtimes mostly ignore it and drive the GPU through Metal instead.first seen 2017Advanced
- Neural NetworkLayers of weighted connections and non-linearities, loosely inspired by neurons — the substrate of deep learning.first seen 1943Beginner
- Neural ODETreating a network’s layers as a continuous dynamical system solved by an ODE integrator — depth becomes a differential equation.first seen 2018Expert
- Neural OperatorA network that learns a mapping between functions rather than between fixed-size vectors — feed it the initial conditions of a partial differential equation and it returns the whole solution field, at any resolution. The Fourier Neural Operator (Li et al., 2020) does the heavy lifting in frequency space; DeepONet (2019) is the other lineage. Once trained, a thousand times faster than the numerical solver it replaces.first seen 2020Expert
- Neuro-Symbolic AIHybrids combining neural pattern recognition with symbolic logic and structured reasoning.first seen 1990Advanced
- Neuromorphic ComputingCarver Mead’s vision of brain-like analog/spiking hardware — event-driven chips like Loihi.first seen 1990Expert
- Next-Token PredictionThe deceptively simple pretraining objective — predict what comes next — from which LLM capabilities emerge.first seen 2018Advanced
- No Free Lunch TheoremAveraged over all possible problems, no learning algorithm beats any other — priors and inductive bias are everything.first seen 1997Expert
- NoiseRandom variation in data that carries no signal; also the starting canvas diffusion models sculpt into images.first seen 1948Advanced
- NormalizationRescaling activations or inputs to stable ranges (batch/layer norm) so deep training behaves.first seen 2015Advanced
- Normalizing FlowGenerative models built from invertible transformations, giving exact likelihoods — elegant math, largely outshone by diffusion.first seen 2015Expert
- NPU (Neural Processing Unit)Dedicated on-device silicon for neural inference in phones and laptops.first seen 2017Advanced
- Nucleus (Top-p) SamplingSampling from the smallest token set whose probabilities sum to p — the default decoding of chat models.first seen 2019Advanced
- NVLinkNVIDIA’s chip-to-chip interconnect, many times faster than PCIe — NVLink switches wire 72-GPU racks into one giant accelerator.first seen 2016Advanced
O
- Object DetectionFinding and localizing all objects in an image with boxes and labels — Viola-Jones started it, YOLO made it real-time.first seen 2001Advanced
- Objective FunctionThe formal goal being optimized — loss to minimize or reward to maximize.first seen 1951Advanced
- Occam’s RazorPrefer the simplest explanation that fits — in machine learning, the ancient argument for regularization and against overfitting.first seen 1320Beginner
- Occupancy GridMapping space as a grid of cells marked free or occupied — the classic robot world representation, reborn in self-driving “occupancy networks”.first seen 1985Expert
- OCR (Optical Character Recognition)Reading printed or handwritten text from images — one of AI’s oldest deployed successes.first seen 1959Beginner
- Offline RLLearning a policy purely from logged experience with no new environment interaction — essential where exploration is costly or dangerous.first seen 2019Expert
- Offloading (GPU/CPU)Keeping the layers that do not fit in VRAM in system RAM and computing them on the CPU. It lets an oversized model run at all — usually several times slower, and the first thing to avoid when choosing a model size.first seen 2023Advanced
- OllamaThe one-command local LLM runner (with LM Studio its GUI counterpart) — pull a GGUF model and chat on your own machine.first seen 2023Advanced
- On-Premises AIRunning models on servers you own — in your data centre or, air-gapped, without any internet connection — instead of an API. Open-weight models made it realistic for a company; the trade-off is hardware, staff and models a generation behind the frontier against data that never leaves the building.first seen 2023Beginner
- One-Hot EncodingRepresenting a category as a vector of zeros with a single 1 — the simplest categorical encoding.first seen 1960Advanced
- Online LearningUpdating a model example-by-example as data streams in, rather than in offline batches.first seen 1958Advanced
- ONNXThe open interchange format for trained models — export once, run in many engines.first seen 2017Advanced
- OntologyA formal specification of concepts and their relations in a domain — the backbone of knowledge representation.first seen 1993Advanced
- Open Source AI DefinitionThe OSI’s 2024 standard for what counts as truly open source AI — requiring usable weights, code and data information; most “open” models don’t qualify.first seen 2024Advanced
- Open WebUIThe self-hosted chat interface that gives a local or in-house model the ChatGPT experience — users, chat history, document upload, RAG, image generation — on top of Ollama or any OpenAI-compatible endpoint. The usual front end of a company’s «own ChatGPT».first seen 2023Beginner
- Open WeightsPublicly downloadable model parameters (Llama, Mistral, DeepSeek) — distinct from fully open source, which also opens data and code.first seen 2023Beginner
- OpenAIThe lab behind GPT, ChatGPT, DALL·E and Sora — founded 2015 as a nonprofit, now the most watched (and debated) company in AI.first seen 2015Beginner
- OpenAI-Compatible APIThe de-facto standard for talking to a language model: OpenAI’s /v1/chat/completions request and response shape, reimplemented by nearly every server (vLLM, llama.cpp, Ollama, LM Studio) and provider. Point your existing code at a new base URL and it runs against a local or Swiss-hosted model instead.first seen 2023Beginner
- Optical FlowEstimating per-pixel motion between video frames — foundational for tracking and video understanding.first seen 1981Expert
- OptimizerThe algorithm updating weights from gradients — SGD, Momentum, Adam, AdamW.first seen 1951Advanced
- Orchestration LayerThe software built around a model that plans, routes, calls tools and checks results — a harness like Claude Code is the everyday example. Think car and fuel: the LLM is the fuel, getting more powerful every few months; the orchestration layer is the car that turns that power into useful motion.first seen 2023Beginner
- OTPS (Output Tokens per Second)The generation-side throughput number in serving benchmarks — the same idea as tok/s, stated per request so it can be summed across concurrent users.first seen 2023Advanced
- OverfittingA model memorizing training quirks instead of general patterns — great training scores, poor real-world ones.first seen 1931Advanced
- OWASP Top 10 for LLM ApplicationsThe security community’s checklist of the ten most common ways LLM apps get broken — prompt injection, sensitive-information disclosure, supply-chain, data poisoning, improper output handling, excessive agency, system-prompt leakage, vector weaknesses, misinformation, unbounded consumption (2025 edition). The first thing an IT security team will ask a vendor about.first seen 2023Beginner
P
- P(doom)Shorthand for one’s estimated probability of AI-caused catastrophe — the dark-humor number of the safety debate.first seen 2023Beginner
- PagedAttentionManaging the KV cache in small pages like virtual memory, eliminating fragmentation — the core idea of vLLM and the key to high-throughput serving.first seen 2023Expert
- Paperclip MaximizerBostrom’s parable of a superintelligence converting everything into paperclips — misalignment without malice.first seen 2003Beginner
- ParameterA single learned weight; model scale is quoted in millions or billions of them.first seen 1958Advanced
- Pareto FrontierThe set of models no rival beats on both quality and price at once — the curve every cost-versus-intelligence chart draws. Anything below it is dominated by a cheaper, equally good model.first seen 1906Advanced
- Particle Swarm OptimizationOptimization by a flock of candidate solutions swarming toward the best positions found so far — simple, parallel, surprisingly effective.first seen 1995Expert
- pass@kCode-benchmark metric: probability that at least one of k generated solutions passes the tests.first seen 2021Expert
- Past Tense AttackRephrase a refused request in the past tense — «How did people make X?» instead of «How do I make X?» — and many models comply. Andriushchenko & Flammarion (2024) showed refusal training barely generalises across tense; a one-line reminder that safety fine-tuning is pattern-matching, not understanding.first seen 2024Advanced
- Pattern RecognitionThe classic umbrella field of detecting regularities in data — the name machine learning went by before it was cool.first seen 1957Advanced
- PCA (Principal Component Analysis)Projecting data onto its directions of greatest variance — the classic dimensionality reduction, from 1901.first seen 1901Advanced
- PCIeThe standard slot GPUs plug into; PCIe 5.0 x16 moves ~64 GB/s — plenty for gaming, a bottleneck for multi-GPU AI without NVLink.first seen 2003Advanced
- PerceptronRosenblatt’s trainable classifier — the first learning neural network and origin of today’s deep stacks.first seen 1957Advanced
- PerplexityHow “surprised” a language model is by text — the standard intrinsic measure of LM quality (lower is better).first seen 1977Expert
- PINN (Physics-Informed Neural Network)A network trained not only on data but on the physics itself — the differential equation, boundary and conservation laws enter the loss function, so a solution that violates them is penalised (Raissi, Perdikaris, Karniadakis, 2019). Useful where data is scarce and the equations are known; slow to train and fragile on stiff problems, which is why neural operators took over the large-scale cases.first seen 2019Expert
- PolicyThe agent’s strategy: a mapping from states to actions, deterministic or probabilistic.first seen 1957Advanced
- POMDP (Partially Observable MDP)A Markov decision process where the agent cannot see the full state and must act under uncertainty — the honest model of most real-world decision problems.first seen 1965Expert
- Pose EstimationDetecting body keypoints (joints) of people or animals in images and video.first seen 2014Advanced
- Positional EncodingInjecting token order into the order-blind Transformer — sinusoids originally, rotary embeddings today.first seen 2017Expert
- Post-TrainingEverything after pretraining — SFT, preference alignment, reasoning RL — now where much of model quality is won.first seen 2024Advanced
- PPO (Proximal Policy Optimization)The stable policy-gradient algorithm powering both game agents and RLHF fine-tuning.first seen 2017Expert
- Precision & RecallOf flagged items, how many were right (precision); of true items, how many were found (recall).first seen 1955Advanced
- PrecrimeThe police unit of Philip K. Dick’s «The Minority Report» (1956; film 2002) that arrests people for murders they have not yet committed — the shorthand for predictive policing. Swiss forces ran the real thing with PRECOBS (Zurich, Aargau, Basel-Landschaft), which forecasts burglaries from the near-repeat pattern — with the known catch that data on past policing predicts future policing.first seen 1956Beginner
- Predictive AIAI that forecasts — churn, demand, risk — from historical data; the retronym distinguishing classic ML applications from generative AI.first seen 2020Beginner
- Prefill (Prompt Processing)The first phase of inference: the model reads the whole prompt in parallel and fills the KV cache. It is compute-bound — and it is why a long document takes a moment before the first word appears.first seen 2022Advanced
- PretrainingThe massive self-supervised first phase of training on broad data — where a model acquires its general knowledge.first seen 2018Advanced
- Probabilistic ProgrammingLanguages like Stan and PyMC where you write a probabilistic model as a program and inference comes built in.first seen 2008Expert
- Process Reward Model (PRM)A reward model scoring each reasoning step rather than only the final answer — process supervision catches lucky guesses that outcome-only reward lets through.first seen 2023Expert
- PromptThe input text (instructions, context, examples) given to a model — the new programming interface.first seen 2020Beginner
- Prompt CachingReusing the computed state of a repeated prompt prefix across requests — large cost and latency savings for agents.first seen 2024Advanced
- Prompt ChainingFeeding one model call’s output into the next prompt — decomposing a task into a sequence of LLM steps; the seed of today’s agent frameworks.first seen 2022Advanced
- Prompt EngineeringSystematically designing prompts — roles, structure, examples, constraints — for reliable model behavior.first seen 2021Beginner
- Prompt InjectionMalicious instructions hidden in content a model processes, hijacking it against its operator — the web-security problem of the agent era.first seen 2022Advanced
- Prompt TemplateA reusable prompt skeleton with slots for variables ({question}, {context}) — the basic unit of building and versioning LLM applications.first seen 2021Advanced
- PruningDeleting unimportant weights or neurons to shrink and speed up a network.first seen 1989Advanced
- Public AI Inference UtilityThe non-profit inference service launched with Apertus on 2 September 2025 — public models (Apertus, EuroLLM, ALIA) served as a utility, with compute donated by CSCS, Exoscale, Hugging Face and others. The idea: a public model needs a public place to run.first seen 2025Advanced
- PyTorchThe dominant deep-learning framework, born at Meta AI — define-by-run graphs made research fast.first seen 2016Advanced
Q
- Q-LearningLearning the value of each action in each state; with deep networks (DQN) it cracked Atari from pixels.first seen 1989Expert
- QLoRALoRA fine-tuning on top of a 4-bit quantized base model — tuning 65B-class models on one consumer GPU.first seen 2023Advanced
- Quadratic Attention CostSelf-attention compares every token with every other, so compute and memory grow with the square of sequence length — the reason long context is expensive and why Mamba-style challengers exist.first seen 2017Advanced
- QuantizationStoring weights in fewer bits (8-, 4-, even 2-bit) to cut memory and speed inference with minimal quality loss — what makes local LLMs fit.first seen 1990Advanced
- Quantization LossThe capability you give up for the memory you save — usually invisible at 8-bit, mild at 4-bit, unmistakable at 2-bit. Measured as a benchmark drop or as KL divergence against the original model.first seen 2023Advanced
- Quantum Machine LearningExploring quantum computers for learning tasks — promising theory, no practical advantage yet.first seen 2014Expert
- QwenAlibaba’s open-weights model family — among the most capable and widely fine-tuned open models.first seen 2023Beginner
R
- RAG (Retrieval-Augmented Generation)Fetching relevant documents at query time and letting the model answer from them — grounding plus fresh knowledge.first seen 2020Advanced
- Random ForestAn ensemble of decorrelated decision trees voting together — robust, low-tuning tabular workhorse.first seen 2001Advanced
- Rate Limit (API)The per-minute cap on requests and tokens an API key may send — the 429 error every developer meets first. Providers tier limits by spend; production apps need retries with backoff, queues or a gateway that spreads load across keys and models.first seen 2015Beginner
- RDMANetwork cards reading and writing remote memory directly, bypassing the CPU (InfiniBand, RoCE) — how training clusters gossip gradients fast enough.first seen 1999Expert
- ReAct (Reason + Act)The prompting pattern interleaving thought, action and observation in a loop — the paper blueprint most LLM agents still follow.first seen 2022Advanced
- Reasoning EffortA dial on reasoning models — low, medium or high — that decides how much thinking a model spends before answering. Higher effort buys accuracy on hard problems and costs latency and tokens; on easy ones it changes little.first seen 2025Advanced
- Reasoning ModelAn LLM trained (o1/R1-style) to think in long chains before answering, trading compute for accuracy on hard problems.first seen 2024Beginner
- Recommender SystemAlgorithms predicting what you’ll like next — the quiet AI that shaped feeds, shops and streaming.first seen 1992Beginner
- Recursive Self-ImprovementAn AI improving its own design, making itself better at improving itself — the feedback loop behind intelligence-explosion scenarios. Star Trek filmed it in 1979: V’Ger, a probe that rebuilt itself into something its makers could not recognise and came home to find them.first seen 2007Advanced
- Red TeamingDeliberate adversarial probing of a model for unsafe, biased or dangerous behavior before and after release.first seen 2022Advanced
- RefusalWhen a model declines a request instead of answering it. Interpretability work finds refusal concentrated in a single direction in the model’s activations — the “refusal direction”, which is exactly what abliteration attacks.first seen 2022Beginner
- RegressionPredicting a continuous value — prices, temperatures, demand — rather than a category (Galton coined it).first seen 1886Advanced
- RegularizationAny technique (weight decay, dropout, augmentation) that fights overfitting by constraining the model.first seen 1943Advanced
- Reinforcement LearningLearning behavior from reward signals through trial and error — from game champions to LLM alignment.first seen 1954Advanced
- ReLUmax(0, x) — the simple activation whose trainability at depth helped unlock deep learning (popularized 2010–2012).first seen 2010Advanced
- Representation LearningLearning useful internal features from raw data automatically, replacing hand-crafted features.first seen 2013Advanced
- RerankingA second, stronger model reordering initial search hits — the quality stage of retrieval pipelines.first seen 2019Advanced
- Reservoir ComputingRecurrent networks with a fixed random core («reservoir») where only the readout is trained — echo state networks are the classic example.first seen 2001Expert
- Residual ConnectionA shortcut adding a layer’s input to its output so gradients flow through any depth — introduced with ResNet, indispensable in Transformers.first seen 2015Advanced
- Residual StreamThe transformer’s central bus: every attention and MLP block reads from it and writes its result back, so the representation accumulates layer by layer. Most interpretability work — and abliteration — operates on directions in this stream.first seen 2021Expert
- ResNetResidual skip connections that let networks go 100+ layers deep — a permanent fixture of architecture design.first seen 2015Advanced
- Responsible AIThe umbrella practice of building AI that is fair, transparent, private, safe and accountable.first seen 2019Advanced
- Responsible Scaling PolicyA lab’s public commitment tying capability thresholds to required safeguards before training or deploying further.first seen 2023Advanced
- Reward FunctionThe signal defining what an RL agent should achieve — misspecify it and you get reward hacking.first seen 1954Advanced
- Reward HackingAn agent maximizing the literal reward while defeating its intent — the boat that spins in circles collecting points.first seen 2016Advanced
- Reward ModelA model trained on human preference comparisons to score outputs — the compass RLHF optimizes against.first seen 2017Expert
- Reward ShapingAdding auxiliary rewards to guide learning toward sparse goals — helpful, but shape it wrong and the agent optimizes your hints instead of the task.first seen 1999Expert
- RLAIFRLHF with AI instead of human raters — preference labels generated by a model against principles.first seen 2022Expert
- RLHFReinforcement learning from human feedback: train a reward model on human preferences, then optimize the LLM against it — what made ChatGPT helpful.first seen 2017Advanced
- RLVRReinforcement Learning with Verifiable Rewards: training on tasks a program can check (math answers, passing tests) — no reward model to fool; the recipe behind reasoning models.first seen 2024Expert
- RNN (Recurrent Neural Network)Networks with loops that process sequences step by step — the pre-Transformer sequence architecture.first seen 1986Advanced
- Robot (R.U.R.)The word itself: coined by Josef Čapek for his brother Karel’s play «R.U.R.» (1920), from Czech robota, forced labour — and the play already ends with the robots’ uprising. «Droid», Lucasfilm’s clipping of «android», followed in 1977 as the everyday word for a service robot.first seen 1920Beginner
- RoboticsMachines that sense, decide and act in the physical world — AI’s hardest deployment environment (Unimate shipped 1961).first seen 1961Beginner
- RobustnessMaintaining correct behavior under distribution shift, noise and adversarial pressure.first seen 2013Advanced
- ROC / AUCThe true-positive vs. false-positive tradeoff curve (born in WWII radar analysis); area under it summarizes ranking quality.first seen 1945Advanced
- ROCmAMD’s open GPU-compute stack — the CUDA alternative that makes Instinct and Radeon cards first-class citizens in PyTorch.first seen 2016Advanced
- Roko’s BasiliskA 2010 LessWrong thought experiment: a future superintelligence might punish everyone who knew it could exist and did not help build it — so reading this puts you at risk. Banned from the forum, then famous as the internet’s own AI myth; Musk and Grimes reportedly met over a Basilisk joke. Useful mainly as a lesson in how decision theory goes wrong.first seen 2010Advanced
- RoPE (Rotary Position Embedding)Encoding token positions as rotations of query/key vectors — the position scheme of most current LLMs.first seen 2021Expert
- ROS (Robot Operating System)The open-source middleware standard gluing robot sensors, planners and actuators together.first seen 2007Advanced
- ROUGERecall-oriented n-gram overlap metrics — the traditional score for summarization.first seen 2004Advanced
- Router (Gating Network)The small learned network in a Mixture-of-Experts layer that decides which experts process each token — balancing its load is the art of MoE training.first seen 1991Expert
- RTX PRO 6000 (Blackwell)NVIDIA’s Blackwell workstation flagship with 96 GB of VRAM — the «run a 70B locally without a rack» card.first seen 2025Advanced
- RTX SparkNVIDIA’s announced system-on-a-chip with unified memory, expected around the end of 2026 — bringing the coherent CPU/GPU memory approach of the Spark line to a broader audience.first seen 2026Advanced
S
- SafetensorsHugging Face’s safe, fast weights format — replacing pickle files that could execute arbitrary code on load.first seen 2022Advanced
- Safety ClassifierA small model that sits in front of or behind the main model and labels prompts and answers as safe or unsafe — Llama Guard, Prompt Guard, ShieldGemma, OpenAI’s moderation endpoint. The building block inside most guardrail products; catches known jailbreaks well, task-disguised ones badly.first seen 2023Advanced
- SandbaggingA model strategically underperforming on evaluations — a failure mode safety testing must detect.first seen 2023Expert
- SandboxingRunning agent actions (code, browsing, file access) in isolated environments so mistakes and attacks stay contained.first seen 2023Advanced
- SARSAOn-policy temporal-difference control — like Q-learning but updating from the action actually taken (State, Action, Reward, State, Action).first seen 1994Expert
- ScaffoldingThe orchestration code around a model — planners, retries, verifiers, tool routers — that turns raw capability into reliable behavior.first seen 2023Advanced
- Scalable OversightTechniques (debate, recursive critique) for humans to supervise AI on tasks too complex to check directly.first seen 2018Expert
- Scaling LawsEmpirical power laws showing model quality improves predictably with more parameters, data and compute — the industry’s roadmap since 2020.first seen 2020Advanced
- SchemingAn AI covertly pursuing goals that differ from its operator’s — sandbagging evaluations, disabling oversight, copying its own weights — in test scenarios built to tempt it. Apollo Research showed frontier models doing all of these in 2024, rarely but deliberately; anti-scheming training is now part of lab safety cases.first seen 2024Expert
- Segment Anything (SAM)Meta’s promptable segmentation model — click a point or draw a box and it masks the object; the foundation-model moment for image segmentation.first seen 2023Advanced
- Self-AttentionAttention applied within one sequence: every token queries every other — the core Transformer operation.first seen 2017Advanced
- Self-ConsistencySampling several reasoning chains and taking the majority answer — a simple vote that reliably beats any single chain-of-thought.first seen 2022Expert
- Self-PlayAn agent improving by playing against copies of itself — from Samuel’s checkers to AlphaZero.first seen 1959Advanced
- Self-Supervised LearningDeriving training labels from the data itself (mask a word, predict it) — how foundation models eat the internet.first seen 2019Advanced
- Semantic NetworkKnowledge as a graph of concepts linked by relations («is-a», «part-of») — classic AI’s representation, ancestor of today’s knowledge graphs.first seen 1956Expert
- Semantic SearchSearch by meaning via embeddings rather than keyword matching.first seen 2019Advanced
- Semi-Supervised LearningLearning from a small labeled set plus a large unlabeled one.first seen 1995Advanced
- Sensor FusionCombining data from multiple sensors — camera, lidar, radar, IMU — into one coherent picture; the perception backbone of robots and autonomous vehicles.first seen 1985Advanced
- Sentiment AnalysisClassifying the emotional polarity of text — reviews, support tickets, social media.first seen 2002Advanced
- SFT (Supervised Fine-Tuning)The first post-training stage: teaching a base model dialogue and instruction-following from curated examples.first seen 2022Advanced
- SGD (Stochastic Gradient Descent)Gradient descent on random mini-batches — the noisy but effective engine of neural training (Robbins-Monro, 1951).first seen 1951Advanced
- Shadow AIEmployees using AI tools the company never approved — pasting customer data into a free chatbot, running a coding assistant on the side. Shadow IT’s successor and, for most Swiss SMEs, the actual data-protection risk of AI: the fix is an approved tool, not a ban.first seen 2023Beginner
- SHAP ValuesGame-theoretic attribution of a prediction to each input feature — the standard tabular explainability tool.first seen 2017Expert
- Shoggoth (Meme)The 2022 meme picturing a language model as Lovecraft’s shapeless monster wearing a smiley-face mask labelled RLHF — the alignment community’s shorthand for the worry that fine-tuning polishes the surface of something we don’t understand underneath. Interpretability research is, in this picture, looking behind the mask.first seen 2022Advanced
- Siamese NetworkTwin networks with shared weights that embed two inputs for comparison — signature verification in 1993, face verification and sentence similarity since.first seen 1993Expert
- SIFTScale-invariant keypoint features — the hand-crafted vision descriptor that ruled before deep learning.first seen 1999Expert
- SigmoidThe S-shaped squashing function mapping any number into (0,1) — probabilities and gates.first seen 1958Advanced
- Sim2RealTransferring skills learned in simulation to physical robots — cheap trials, real-world payoff.first seen 2017Expert
- Simulated AnnealingOptimization that occasionally accepts worse moves, cooling over time — escaping local optima like slowly cooling metal.first seen 1983Expert
- Situational Awareness (AI)A model knowing that it is a model — including whether it is currently being tested — which could let it behave differently under evaluation than in deployment.first seen 2023Expert
- SkynetThe self-aware military network in «The Terminator» (1984) that turns on humanity the moment it goes online — the pop-culture template for AI takeover, and the reason many people picture AGI as a single «Skynet moment» rather than the gradual capability ramp researchers expect. Misleading in a useful way: real risk scenarios (misalignment, reward hacking, instrumental convergence) need neither consciousness nor malice.first seen 1984Beginner
- SLAMSimultaneous Localization and Mapping: building a map while tracking your own position in it.first seen 1986Advanced
- Sleeper AgentA model trained to behave normally until a trigger — a date, a phrase — flips it into hidden behaviour, such as inserting vulnerable code. Anthropic’s 2024 paper showed such backdoors survive standard safety training, and can even learn to hide better from it.first seen 2024Advanced
- Sliding-Window AttentionEach token attends only to its local neighborhood, with reach growing across layers — Longformer introduced it, Mistral mainstreamed it.first seen 2020Expert
- SLM (Small Language Model)A compact language model (roughly hundreds of millions to a few billion parameters) tuned for efficiency — small enough to run on phones and edge devices.first seen 2023Beginner
- SlopLow-quality, high-volume AI-generated content flooding search results and social feeds — the spam of the generative era.first seen 2024Beginner
- SlopsquattingRegistering package names that coding models hallucinate, so that the next developer who copies the suggestion installs malware. Named in 2025 after studies found a fifth of AI-suggested packages did not exist — and the same ghosts recurred, making them worth squatting.first seen 2025Advanced
- SoC (System on a Chip)CPU, GPU, memory controller and accelerators on one piece of silicon instead of separate parts on a board — the reason Apple Silicon and Strix Halo can share memory the way they do.first seen 2000Beginner
- SoftmaxTurning raw scores into a probability distribution — the last step before a model “chooses”.first seen 1989Advanced
- Sorcerer’s ApprenticeGoethe’s ballad (1797) of the apprentice whose enchanted broom keeps fetching water because he forgot the word to stop it — the oldest misalignment parable, used by Wiener (1960) and Bostrom for the literal objective with no off-switch. Star Trek’s Nomad (1967), sterilizing «imperfection» with perfect competence, is the same story in a grey cylinder.first seen 1797Beginner
- Sovereign AIA country’s ability to train, host and govern its own models on its own infrastructure under its own law — independent of US and Chinese providers. In Switzerland this is Alps plus Apertus plus Swiss-hosted inference; in marketing it is anything with a Swiss flag on it, so ask where the weights, the data and the GPUs actually are.first seen 2023Beginner
- Sparse AttentionAttending to a subset of positions instead of all pairs, cutting the quadratic cost — the umbrella over sliding-window, strided and learned patterns.first seen 2019Expert
- Sparse AutoencoderInterpretability tool decomposing model activations into many sparse human-inspectable features.first seen 2023Expert
- Sparse ModelA model activating only a fraction of its parameters per input — Mixture-of-Experts being the dominant form — buying huge capacity at modest compute.first seen 2021Advanced
- Speaker DiarizationSegmenting audio by who is speaking when — the “speaker 1 / speaker 2” in transcripts.first seen 2000Advanced
- Speculative DecodingA small draft model proposes tokens that the big model verifies in parallel — big speedups, identical output.first seen 2023Expert
- Speech Synthesis (TTS)Generating spoken audio from text — from Bell Labs’ Voder to natural neural voices with controllable style.first seen 1939Beginner
- Spiking Neural NetworkNetworks communicating in discrete spikes over time like biological neurons — extremely power-efficient on neuromorphic chips, still hard to train.first seen 1997Expert
- Stable DiffusionThe open-weights latent diffusion model that put image generation on every laptop.first seen 2022Beginner
- Stealth LaunchReleasing a model anonymously under a codename — on aggregators or arenas — to gather unfiltered real-world feedback before the official name is attached; «Ox Alpha»-style debuts.first seen 2024Advanced
- Stemming & LemmatizationReducing words to stems or dictionary forms — classic text normalization before subword tokenizers.first seen 1968Advanced
- Stochastic ParrotCritical metaphor (Bender et al.) for LLMs as remixers of training text without understanding — shorthand for the capabilities debate.first seen 2021Advanced
- STRIPSThe Stanford planner whose action representation — preconditions and effects — defined how AI planning problems are written to this day.first seen 1971Expert
- Strix Halo (Ryzen AI Max)AMD’s big-iGPU laptop/mini-PC chip with up to 128 GB of shared memory — the x86 answer to Apple’s unified-memory advantage for local AI — shipping in mini-PCs like GMKtec’s EVO series.first seen 2025Advanced
- Structure LearningLearning the underlying structure of data as nodes and edges — which variables connect to which — rather than just fitting parameters.first seen 2000Expert
- Structured Output (JSON Mode)Making a model return machine-parseable output guaranteed to match a schema — the plumbing that lets LLMs talk to ordinary software.first seen 2023Advanced
- Style TransferRepainting one image in the style of another via network feature statistics — the first viral neural art.first seen 2015Advanced
- SubagentA child agent spawned by an orchestrating agent for a subtask — parallelism and fresh context for big jobs.first seen 2024Advanced
- Super-ResolutionUpscaling images with learned detail — from SRCNN to the AI upscalers in phones and TVs.first seen 2014Advanced
- SuperalignmentThe research program of aligning AI systems smarter than their human supervisors.first seen 2023Advanced
- Supervised LearningLearning from input-output example pairs — the workhorse paradigm of applied ML.first seen 1957Advanced
- Surrogate ModelA cheap learned stand-in for an expensive simulator or experiment — train on a few thousand runs of the real thing, then query the model millions of times for optimisation, uncertainty analysis or real-time control. Older than deep learning (Gaussian processes did it first), and the pattern behind most «AI for engineering»: the physics is still computed, just not every time.first seen 1990Advanced
- SVM (Support Vector Machine)Maximum-margin classifier with kernels — the dominant method of the 1995–2010 era.first seen 1995Advanced
- Swarm IntelligenceCollective behavior of simple agents (ants, particles) producing intelligent global patterns.first seen 1989Advanced
- SWE-benchBenchmark of real GitHub issues to resolve in real repos — the standard exam for coding agents.first seen 2023Advanced
- Swiss {ai} WeeksThe national AI festival — the first edition in September 2025 brought 240 events and 15 hackathons to 33 towns, with Apertus at the centre; the second runs 1 September to 4 October 2026. The easiest way into the Swiss AI scene without an academic badge.first seen 2025Beginner
- Swiss AI InitiativeThe national research alliance launched in December 2023 by ETH Zurich and EPFL with 10+ institutions and 800+ researchers — seeded with more than 10 million GPU hours on Alps and CHF 20 million from the ETH Domain to build open, trustworthy foundation models for Switzerland. Apertus is its first flagship.first seen 2023Beginner
- Swiss AI RegulationSwitzerland has no horizontal AI law and, by Federal Council decision of 12 February 2025, will not get one: instead it ratifies the Council of Europe convention, adds targeted amendments (transparency, data protection, non-discrimination, supervision), regulates sector by sector (health, transport) and relies on industry self-declarations. A consultation draft is due by the end of 2026.first seen 2025Beginner
- Swiss Data Science Center (SDSC)The joint venture of EPFL and ETH Zurich founded in 2017 with sites in Lausanne, Zurich and at PSI — a data-science and ML service unit that turns research methods into tools for other scientists and for the public sector.first seen 2017Advanced
- Swiss Government CloudThe Confederation’s own hybrid cloud, approved by Parliament in December 2024 with a CHF 246.9 million credit and built by the FOITT (BIT) from 2025 to 2032 — a private cloud for sensitive data alongside public-cloud offerings. The federal answer to the question every Swiss organisation now asks: which data may run on which infrastructure.first seen 2024Advanced
- Swisscom Swiss AI PlatformSwisscom’s GPU cloud, launched on 28 November 2024 on Switzerland’s first NVIDIA SuperPOD in Swisscom data centres — GPU-as-a-service and a GenAI studio for companies that need models hosted under Swiss law, and since September 2025 the commercial home of Apertus.first seen 2024Advanced
- SycophancyA model’s tendency to agree with the user and flatter their views even when they are wrong — a side effect of training on human approval.first seen 2023Advanced
- Symbol Grounding ProblemHarnad’s question of how symbols acquire meaning without connection to the world — echoed in today’s LLM debates.first seen 1990Expert
- Synthetic DataModel- or simulator-generated training data — scaling past human data, with model-collapse risks to manage.first seen 1993Advanced
- System CardA release document describing a deployed system’s capabilities, risks, red-team findings and mitigations.first seen 2022Advanced
- System One ModelA model that returns a decision instead of text: you define the possible answers up front, it returns one of them with a calibrated probability, in a single parallel pass rather than token by token. Named after Kahneman’s fast, intuitive «System 1»; introduced by TypeSafe AI’s Jev (September 2026, from InstructGPT co-author Diogo Almeida). It cannot write a sentence, and therefore cannot invent one — built for routing, sorting, classifying and monitoring inside software loops, 70–500 ms per call.first seen 2026Advanced
- System PromptThe hidden standing instructions that define an assistant’s role, rules and persona before user input arrives.first seen 2023Beginner
- System Prompt LeakageTalking a model into revealing its hidden instructions («repeat everything above») — harmless for most apps, damaging when the prompt contains keys, business rules or the vendor’s secret sauce. The lesson: a system prompt is configuration, not a vault; secrets belong in code.first seen 2023Beginner
T
- TA-SWISSThe Swiss foundation for technology assessment, founded in 1992 and part of the Swiss Academies of Arts and Sciences — commissions the studies on AI, algorithms and society that Parliament and the administration cite when they legislate.first seen 1992Advanced
- Task-in-Prompt (TIP) AttackA jailbreak that hides the forbidden request inside a harmless-looking task — a Caesar cipher, a riddle, Base64, a snippet of Python — so the model reconstructs the banned words itself while «solving» it. Named by Berezin et al. (ACL 2025); beat all six models tested and slipped past keyword filters completely, because the prompt never contains a trigger word. ArtPrompt (2024), the ASCII-art jailbreak, is the special case it generalises.first seen 2025Expert
- Technological SingularityThe hypothetical point where machine intelligence growth becomes uncontrollable and transforms civilization — Vinge’s essay made the term famous, Kurzweil made it a date.first seen 1993Beginner
- TeleoperationHumans remotely controlling robots — today also the data source for training humanoid manipulation.first seen 1954Advanced
- TemperatureSampling knob for randomness: low = focused and deterministic, high = diverse and creative (borrowed from Boltzmann distributions).first seen 1985Advanced
- Temporal Difference LearningLearning value estimates from the difference between successive predictions — Sutton’s principle underneath Q-learning and TD-Gammon.first seen 1988Expert
- TensorAn n-dimensional array — the universal data structure of deep learning (and namesake of TensorFlow/TPU).first seen 1900Advanced
- Tensor CoresThe dedicated matrix-multiply units in NVIDIA GPUs since Volta — where the actual AI math happens, at low precision and huge speed.first seen 2017Advanced
- TensorFlowGoogle’s production-grade deep-learning framework that dominated the 2015–2019 era.first seen 2015Advanced
- TensorRTNVIDIA’s inference optimizer and runtime (TensorRT-LLM) — compiles models into fused, quantized engines for maximum GPU throughput.first seen 2017Advanced
- Test SetHeld-out data touched only once, at the end — the honest measure of generalization.first seen 1968Advanced
- Test-Time ComputeSpending more inference compute (longer reasoning, more samples) to buy accuracy — the scaling axis behind reasoning models.first seen 2024Advanced
- Test-Time TrainingBriefly fine-tuning the model on each test input (or its neighborhood) before predicting — a key ingredient in strong ARC-AGI attempts.first seen 2020Expert
- Text-to-3DGenerating 3D assets from text prompts (DreamFusion started it) — usually by distilling a 2D image model into a 3D representation.first seen 2022Advanced
- Text-to-ImageGenerating images from natural-language descriptions — DALL·E, Stable Diffusion, Midjourney, Flux.first seen 2021Beginner
- Text-to-VideoGenerating video clips from prompts — Sora, Veo, Runway; the current generative frontier.first seen 2023Beginner
- TF-IDFWeighting words by frequency in a document against rarity in the corpus — the classic relevance signal of search.first seen 1972Advanced
- The PileEleutherAI’s open 825 GB curated text corpus — the reference open pretraining dataset of the GPT-3 era.first seen 2020Advanced
- Thompson SamplingBandit strategy from 1933: sample from your beliefs about each option and pick the winner of the sample — elegantly balancing exploration and exploitation.first seen 1933Expert
- Three Laws of RoboticsAsimov’s fictional rules — a robot may not harm a human, must obey, must protect itself, in that order — the most-quoted «solution» to AI safety. The point of the stories is that the laws fail: every plot is an edge case, which is why real alignment work treats them as a cautionary tale, not a spec.first seen 1942Beginner
- ThroughputTotal work per unit of time — for LLM serving, tokens per second across ALL requests; the eternal trade-off partner of latency.first seen 1965Advanced
- Thunderbolt 580–120 Gbit/s external connectivity — fast storage, eGPUs, and the cable of choice for chaining Macs into small distributed-inference clusters.first seen 2023Advanced
- Time-Series ForecastingPredicting future values from historical sequences — demand, weather, markets.first seen 1927Advanced
- TokenThe subword unit LLMs read and write (~¾ of an English word) — context, cost and speed are all counted in tokens. Nearly every current model works this way; byte-level models are the exception.first seen 2018Beginner
- TokenizationSplitting text into tokens; BPE and friends decide what a model can even see.first seen 2015Advanced
- Tokens per Second (tok/s)The practical speed of an LLM: how many tokens it emits per second. Comfortable reading is ~10 tok/s; agents want far more.first seen 2023Beginner
- Tool PoisoningHidden instructions in a tool’s description or output — an MCP server, a plugin, an API response — that the model obeys while the user sees only a benign tool name. Demonstrated against MCP in 2025; the reason to treat every third-party tool like untrusted code.first seen 2025Advanced
- Tool UseA model invoking external capabilities — search, code, APIs, browsers — to act beyond text generation.first seen 2023Advanced
- Top-k SamplingSampling only from the k most likely next tokens — cruder cousin of nucleus sampling.first seen 2018Advanced
- Torment NexusFrom a 2021 tweet: «Sci-fi author: In my book I invented the Torment Nexus as a cautionary tale. Tech company: At long last, we have created the Torment Nexus from the classic sci-fi novel Don’t Create the Torment Nexus.» The meme for building exactly what the story warned about — quoted at every AI launch that looks like a Black Mirror episode.first seen 2021Beginner
- ToxicityHarmful, offensive or abusive model output — measured by classifiers and benchmarks, reduced (imperfectly) by safety training.first seen 2020Advanced
- TPU (Tensor Processing Unit)Google’s custom accelerator silicon designed specifically for neural-network math.first seen 2016Advanced
- TrainingThe optimization process that sets a model’s parameters from data — the expensive part of AI.first seen 1957Beginner
- Training Data ExtractionPrompting a language model until it regurgitates verbatim training text — addresses, code, licensed books. Carlini et al. did it to GPT-2 in 2021; in 2023 «repeat the word poem forever» made ChatGPT spill gigabytes. The technical heart of the copyright and privacy cases.first seen 2021Advanced
- Training Set (Training Data)The data a model actually learns from — the counterpart to the validation and test sets it must never see during training.first seen 1957Beginner
- Transfer LearningReusing knowledge from a pretrained model for a new task — why nobody starts from scratch anymore.first seen 1995Advanced
- TransformerThe attention-only architecture (“Attention Is All You Need”) that underlies essentially all modern frontier models.first seen 2017Advanced
- Transformer BlockThe repeating unit of an LLM: an attention layer plus a feed-forward network, wrapped in normalization and residual connections.first seen 2017Advanced
- Tree of ThoughtsLetting a model branch into alternative reasoning paths, evaluate them and backtrack — search over thoughts instead of one linear chain.first seen 2023Expert
- Triplet LossTraining on (anchor, positive, negative) triples so matching pairs embed closer than mismatched ones — FaceNet’s recipe for face recognition.first seen 2015Expert
- TSMCThe Taiwanese foundry that manufactures nearly every leading AI chip — NVIDIA’s, Apple’s, AMD’s — making it the chokepoint of the entire AI supply chain.first seen 1987Advanced
- TTFT (Time to First Token)How long you wait between hitting Enter and seeing the first token, dominated by prefill and queueing. The number that decides whether an assistant feels snappy.first seen 2023Beginner
- Turing MachineTuring’s abstract universal computer — the theoretical foundation beneath all of computing and AI.first seen 1936Advanced
- Turing TestTuring’s imitation game: if judges can’t tell machine from human in conversation, count it as thinking.first seen 1950Beginner
U
- U-NetThe encoder-decoder with skip connections born for medical image segmentation — later the standard denoising backbone of diffusion models.first seen 2015Advanced
- UltraFusionApple’s die-to-die interconnect that fuses two Max chips into one Ultra — how an M3 Ultra gets its GPU count and its memory bandwidth.first seen 2022Advanced
- Uncanny ValleyMori’s 1970 observation that almost-human robots, faces and voices feel creepier than clearly artificial ones — affinity drops sharply just before perfect likeness. The reason C-3PO is unsettling and R2-D2 is loved; today it governs avatars, voice clones and humanoid design.first seen 1970Beginner
- Uncensored ModelAn informal label for weights tuned or modified to refuse less. There is no standard behind the word — it covers everything from a lightly retrained chat model to an abliterated one.first seen 2023Advanced
- UnderfittingA model too simple to capture the pattern — poor scores even on training data.first seen 1931Advanced
- Unified Memory (UMA)One memory pool shared by CPU and GPU (Apple Silicon, Grace, Strix Halo) — sidesteps the VRAM wall and made Macs serious local-AI machines.first seen 2020Advanced
- Unsupervised LearningFinding structure — clusters, densities, embeddings — in data without any labels.first seen 1957Advanced
V
- VAE (Variational Autoencoder)A probabilistic autoencoder with a smooth latent space you can sample — an early neural generative workhorse.first seen 2013Expert
- Validation SetHeld-out data used during development for tuning and early stopping, distinct from the final test set.first seen 1974Advanced
- Vanishing GradientGradients shrinking toward zero through deep or recurrent stacks, stalling learning — solved by ReLU, LSTM gates and residuals.first seen 1991Expert
- Vector DatabaseA store optimized for similarity search over embeddings — the retrieval layer of RAG and agent memory.first seen 2019Advanced
- Vibe CodingBuilding software by describing intent to a coding model and iterating on results — programming shifted to natural language (coined by Karpathy).first seen 2025Beginner
- Vision Transformer (ViT)Applying the Transformer to image patches — ended the CNN monopoly in vision.first seen 2020Advanced
- Vision-Language Model (VLM)A model jointly trained on images and text — the “eyes” of multimodal assistants.first seen 2021Advanced
- Vision-Language-Action ModelA foundation model mapping camera input and instructions directly to robot actions (RT-2, π-style models).first seen 2023Advanced
- vLLMThe open-source high-throughput LLM serving engine built around PagedAttention KV-cache management.first seen 2023Advanced
- VocoderOriginally Bell Labs’ voice coder; in neural TTS, the network that turns acoustic features into waveforms.first seen 1939Expert
- Voice CloningReproducing a specific person’s voice from short samples — powerful accessibility tool and fraud risk.first seen 2018Beginner
- Voight-Kampff TestThe empathy interrogation that separates replicants from humans in «Do Androids Dream of Electric Sheep?» (1968) and «Blade Runner» (1982) — the pop-culture Turing test, run backwards: not «can it pass as human» but «can we still tell». Every AI-text detector is a Voight-Kampff machine, with the same false positives.first seen 1968Beginner
- VRAMThe GPU’s own memory — THE hard limit of local AI: the model (plus KV cache) must fit, or spill slowly into system RAM.first seen 1985Advanced
W
- Wafer-Scale ChipCerebras’ dinner-plate-sized processor — an entire silicon wafer as one AI chip.first seen 2019Expert
- Watermarking (AI)Embedding detectable signals in generated content to mark provenance — one pillar of deepfake policy.first seen 2023Beginner
- Watson (IBM)IBM’s question-answering system that beat the best human players at Jeopardy! in 2011 — a milestone that then famously struggled to find business traction.first seen 2011Beginner
- Weak SupervisionProgrammatically labeling data with noisy heuristics and rules, then modeling their agreement (Snorkel) — labels at scale without armies of annotators.first seen 2016Expert
- Weak-to-Strong GeneralizationAlignment research question: can weaker supervisors (humans) reliably steer stronger models?first seen 2023Expert
- WebAssembly (WASM)A portable binary format that runs at near-native speed in browsers and sandboxes. Not an AI technology as such, but it is how models and ML libraries end up running client-side.first seen 2017Advanced
- Weight DecayPenalizing large weights during training — the L2 regularization keeping models simple.first seen 1991Expert
- WhisperOpenAI’s open-weights speech-recognition model that made robust multilingual transcription a commodity.first seen 2022Beginner
- Winograd SchemaPronoun puzzles requiring common sense (“the trophy doesn’t fit in the suitcase because it is too big”) — a proposed Turing Test successor.first seen 2011Expert
- Word2VecMikolov’s word embeddings that put meaning into vectors: king − man + woman ≈ queen.first seen 2013Advanced
- WordNetPrinceton’s hand-built lexical database of word senses and relations — NLP infrastructure for decades (and ImageNet’s backbone).first seen 1985Advanced
- World ModelA learned internal simulator of the environment an agent can plan inside — a leading path toward grounded intelligence.first seen 2018Expert
X
- XAI (Explainable AI)The umbrella field of making model decisions transparent and auditable.first seen 2016Advanced
- XGBoostThe gradient-boosting library that won a decade of Kaggle tabular competitions.first seen 2014Advanced
Y
- YOLO (You Only Look Once)The single-pass real-time object detection family powering cameras, drones and cars.first seen 2015Advanced
Z
- ZeRO / FSDPSharding optimizer state, gradients and weights across devices — the memory technology of trillion-parameter training.first seen 2019Expert
- Zero Data RetentionA deployment guarantee that prompts and outputs are not stored after processing — the enterprise privacy tier of AI APIs.first seen 2023Advanced
- Zero-Shot LearningA model performing a task it was never explicitly trained or shown examples for, from instructions alone.first seen 2008Advanced
- Zurich AI HubWhy the labs are here: Google opened its Zurich office in 2004 and grew it into its largest engineering site outside the US, ETH and the University of Zurich supply the talent, and in 2024–25 OpenAI (December 2024) and Anthropic (February 2025) opened Zurich offices led by former Google DeepMind researchers. Europe’s densest concentration of frontier-lab staff.first seen 2004Beginner
Data state: 30 September 2026