How AI Systems Are Evaluated: Methods, Benchmarks, and Human Review

How AI systems are evaluated has become one of the most contested questions in technology, because every method used to grade these systems has a documented flaw. Benchmarks saturate and leak into training data, human voters are noisy and gameable, AI judges favor their own outputs, and production failures regularly appear in models that aced every exam. This explainer maps the full evaluation stack with named, verified evidence for each layer.
AI evaluation methods now form a stack of five layers, each answering a different question. Benchmarks ask "what does the model know?" through standardized tests with automatic scoring. Human review asks "which output do people actually prefer?" through blind votes and expert ratings. LLM-as-judge asks "can another AI grade this at scale?" Safety evaluations and red teaming ask "what could this model do if misused?" before release. And production monitoring asks "is it still working?" after real users arrive. Understanding how each layer works - and how each one fails - explains why two models with identical benchmark scores can behave completely differently in your hands.
Quick answer: AI systems are evaluated through (1) benchmarks - standardized tests like MMLU and SWE-bench with automatic scoring, which saturate and leak into training data; (2) human review - blind preference votes like Chatbot Arena and expert annotation, which are noisy and gameable; (3) LLM-as-judge - AI models grading AI outputs at scale, roughly 80-85% aligned with humans but carrying position, verbosity and self-preference biases; (4) safety evals and red teaming - adversarial testing that now gates releases and is required for systemic-risk models under the EU AI Act; and (5) production monitoring - drift detection, RAG metrics and A/B tests after deployment. No single layer is trustworthy alone.
Table of Contents
- The short answer
- How I mapped the evaluation stack: three tests
- Benchmarks: the standardized exam layer
- Human review: the preference layer
- LLM-as-judge: the scale layer
- Safety evals and red teaming: the deployment gate
- Regulation: when evaluation becomes law
- Production evaluation: where models meet reality
- The strongest counterargument, answered
- What to actually do before shipping AI
- FAQs
- The Bottom Line
The short answer
AI systems are evaluated by a layered stack, not a single test. Automated benchmarks (MMLU, SWE-bench, GPQA, Humanity's Last Exam) score knowledge and skills with fixed question sets; human preference systems (Chatbot Arena, expert panels) rank models by blind votes; LLM judges grade millions of outputs at roughly 80-85% agreement with humans; safety evaluations and red teaming probe for misuse risks before release; and production monitoring tracks drift, faithfulness and user outcomes after deployment. Every layer carries documented weaknesses - saturation, vote-rigging, judge bias, voluntary enforcement, and benchmark-to-production gaps - which is why serious evaluation stacks the layers instead of trusting any one of them.
How I mapped the evaluation stack: three tests
I applied three tests to decide what belongs in this explainer. First, the primary-source test: every load-bearing claim needed a named paper, official report, statute text or system card behind it - vendor marketing alone did not qualify. Second, the 2026-relevance test: each method had to be actively shaping how models are graded this year, not retired history. Third, the failure-evidence test: for every layer, I looked for documented cases where the method produced a wrong or misleading verdict, because an evaluation method is best understood through its failure modes.
Where numbers came from aggregators I could not verify against primary sources, they are either omitted or explicitly flagged as unverified below. Two commonly repeated claims failed this test outright and were excluded: precise enterprise "evaluation maturity" percentages (no rigorous survey exists) and unverified leaderboard snapshots citing model names that no official page confirms. What remains is the layer stack that labs, regulators and enterprises actually use in 2026.
Benchmarks: the standardized exam layer
Benchmarks are fixed question sets with automatic scoring, run identically across models to produce comparable numbers. The main families: knowledge (MMLU and its harder successor MMLU-Pro), code (HumanEval, then SWE-bench and SWE-bench Verified, which require resolving real GitHub issues), reasoning (GPQA Diamond's graduate-level science questions and ARC-AGI's novel-puzzle tests), and expert composite exams - most famously Humanity's Last Exam, a Nature-published benchmark of roughly 2,500 expert-written questions built specifically because earlier tests had saturated.
Saturation is the layer's defining disease. The Stanford HAI 2026 AI Index reports that on SWE-bench Verified, performance rose from 60% to near 100% in a single year - meaning a benchmark that took years to build stopped discriminating between frontier models within months. MMLU crossed 90% among frontier models, and on GPQA Diamond the top models cluster between roughly 93% and 94%, a spread smaller than the benchmark's own measurement noise - statistically meaningless ranking territory. Humanity's Last Exam bought the field fresh headroom, and frontier models gained about 30 percentage points on it within a single year per the same AI Index.

What are AI benchmarks and why do they saturate?
A benchmark is a standardized test with a fixed set of questions and automatic scoring - the AI equivalent of a sit-down exam. Benchmarks saturate because models improve faster than tests are rebuilt: once most frontier models score above 90%, the remaining differences fall inside measurement noise, so the benchmark can no longer separate good from great. Contamination accelerates the slide - Benchmark Data Contamination is the documented phenomenon of test questions leaking into training data, producing artificially high scores - which is why fresh-question benchmarks like LiveBench release new items monthly. Research in 2026 found even dynamic benchmarks "do reduce contamination, they do not eliminate it."
Human review: the preference layer
Human evaluation answers the question benchmarks cannot: which output do people actually prefer? The most influential implementation is Chatbot Arena (now LMArena), where users type their own prompts, receive two anonymous model responses side by side, and vote for the better one. Millions of these blind pairwise votes feed a Bradley-Terry statistical model that produces the Elo-style leaderboard the industry watches. The platform added "style control" in August 2024 to statistically separate substance from formatting - a direct response to evidence that voters reward length and markdown polish as much as correctness.
The layer's problems are equally documented. A 2025 paper titled "The Leaderboard Illusion" (arXiv:2504.20879) alleged that private testing of multiple model variants before public release, selective disclosure of scores, and unequal data access favored large labs - charges the Arena operators disputed but which triggered lasting skepticism. An ICML 2025 study demonstrated that the votes themselves are "noisy and biased" and that a few hundred rigged votes could materially shift rankings. And the Llama 4 Maverick episode in April 2025 - where a model listed on the Arena behaved differently from the released version - showed even the identity of what you are voting on can be contested. Underneath all of it sits the oldest problem: human annotator agreement is "notoriously unstable," as Chiang and Lee documented at ACL 2023.

How does Chatbot Arena ranking work?
Users submit a prompt, two anonymous models answer, and the user votes for the better response without knowing which is which. Those millions of pairwise votes are converted into ratings with a Bradley-Terry statistical model (the modern replacement for raw Elo), with bootstrap confidence intervals and a "style control" adjustment that discounts length and formatting effects introduced in August 2024. The design is clever - anonymity prevents brand bias, volume drowns noise - but research has documented vote-rigging potential, preferential data access for big labs, and style biases, so rankings should be read as one signal among several rather than ground truth.
LLM-as-judge: the scale layer
Human votes are expensive, so the field automated them: use a strong AI model - typically a frontier LLM - to score, rank or pass/fail the outputs of other AI systems. The canonical evidence comes from the MT-Bench paper (Zheng et al., NeurIPS 2023), which showed GPT-4 acting as judge matched a panel of human experts on roughly 80% of comparisons - and 85% on non-tie votes - a level comparable to agreement between the humans themselves. The economic driver is blunt: a judge costs "orders of magnitude" less than a human annotator (arXiv, July 2026), which is why CI pipelines and vendor dashboards now grade most AI output with another AI.
The documented costs are three systematic biases. Position bias: verdicts flip depending on which answer is presented first, so practitioners now run the comparison twice with swapped order and average. Verbosity bias: measured length-aware, models like Llama and Gemini variants show classic longer-is-better effects in the +0.24 to +0.44 range (OpenReview). And self-preference: judges score their own family's outputs higher even when human annotators rate them equal (Panickssery et al., arXiv:2404.13076) - a structural conflict of interest when a lab's flagship model grades its own siblings. Mitigations exist - swapped-order averaging, explicit rubrics and checklists (which 2025 research showed can beat larger fine-tuned judges), post-hoc calibration, and open specialist judges like Prometheus 2 - but a July 2026 identifiability paper argues some biases cannot be fully corrected after the fact. The practitioner heuristic that emerged: treat a judge as trustworthy only above roughly 80% measured agreement with humans on your specific task.

Is LLM-as-a-judge reliable?
Reliable within known limits. The MT-Bench research line measured 80-85% agreement with human experts on well-curated general chat questions, and a June 2026 validation study covering 21 judges from nine providers is the largest systematic check to date. But reliability degrades on niche domains, degrades with weak judge models, and the three documented biases - position, verbosity, self-preference - mean raw judge scores are systematically distorted rather than neutral. With swap-order averaging, explicit rubrics, and a different model family doing the judging, it becomes a dependable scale layer; without them, it silently reshapes whatever it measures.
Safety evals and red teaming: the deployment gate
Benchmarks measure capability; safety evaluations measure what happens when capability is pointed at the wrong target. Red teaming is the adversarial probe: "simulate adversarial behavior... to find gaps in how an AI model performs, responds, or fails." It runs at three scales - internal dedicated teams, contracted external experts (OpenAI formalized its external network in a March 2025 paper), and public bug bounties: Anthropic pays up to $15,000 for novel universal jailbreaks through its HackerOne-hosted program. Every frontier lab now pairs these probes with a published safety framework that defines capability thresholds and mandated evals.
The frameworks differ in name but share a skeleton. OpenAI's Preparedness Framework (December 2023) tracks CBRN, cybersecurity and AI-self-improvement risks against HIGH/CRITICAL scorecards. Anthropic's Responsible Scaling Policy defines ASL (AI Safety Level) thresholds with mandated capability evals at least every six months - and in May 2025, Claude Opus 4 became the first model to trigger ASL-3 activation, deploying enhanced CBRN-focused safeguards, with a designated Responsible Scaling Officer empowered to pause training or deployment. Google DeepMind's Frontier Safety Framework evaluates models against Critical Capability Levels "before deployment," adding a harmful-manipulation level in September 2025. Independent testing has entered the loop too: OpenAI's GPT-5 system card confirms pre- and post-deployment evaluations with both the US CAISI and the UK AI Security Institute.

Does any of this actually gate releases? The strongest 2026 evidence says yes: in late September 2026, OpenAI publicly said it would not release its GPT-6.1 Astra model after safety testing found it failed to meet the company's standards - the first headline case of evals visibly stopping a flagship launch (reported by multiple outlets; OpenAI's statement is the primary source). The counter-evidence is equally real: these frameworks are self-authored, Anthropic removed its hard training limit in February 2026, and an independent April 2026 assessment scored DeepMind's framework at 10% on specificity - meaning the gate is real but the gatekeeper is still the model's own maker.
What is red teaming for AI models?
Red teaming is structured adversarial testing: internal teams, hired external experts, or paid bounty hunters deliberately try to make a model misbehave - produce dangerous instructions, leak training data, bypass refusals, or execute harmful agent actions. Frontier labs run all three tiers simultaneously, and the results feed safety frameworks that define deployment thresholds. It is point-in-time testing by design, which means it certifies the model that was tested, not the system users will meet after updates and real-world drift.
Regulation: when evaluation becomes law
In 2026, some evaluation stopped being voluntary. The EU AI Act's general-purpose AI obligations took effect on August 2, 2025: Article 53 requires every GPAI provider to maintain technical documentation and information for downstream developers, while Article 55 requires providers of models classified as systemic-risk - presumed when training exceeds 10^25 FLOP - to "perform model evaluation in accordance with standardised protocols," conduct adversarial testing, and report serious incidents. The July 2026 Digital Omnibus deferred the Act's high-risk-system obligations to December 2027, but explicitly did not defer the GPAI model-level duties - model evaluation documentation is law now, not a proposal. The US approach runs voluntary by comparison: NIST's AI Risk Management Framework and its Generative AI Profile (2024) define a Govern/Map/Measure/Manage cycle that federal agencies and vendors treat as the baseline, without statutory force.

Are AI evaluations required by law?
In the EU, yes for general-purpose AI models: since August 2, 2025, Article 53 requires technical documentation from all GPAI providers, and Article 55 adds standardized-protocol model evaluation, adversarial (red-team) testing and incident reporting for models presumed systemic-risk above 10^25 FLOP of training compute. The Digital Omnibus of July 2026 postponed the separate high-risk-system rules to December 2027 but left these GPAI duties in force. In the US, the NIST AI Risk Management Framework remains voluntary guidance, though a 2026 federal framework moves toward mandatory pre-release testing of the most powerful models.
Production evaluation: where models meet reality
The final layer begins after launch, and it is where the benchmark story most often falls apart. A February 2026 arXiv paper states the structural problem plainly: evaluation pipelines "themselves introduce silent failure modes" and "high benchmark scores routinely mislead teams." Production evaluation therefore measures a different object - not the model, but the system: retrieval plus generation plus prompts plus users. The distinction matters because a system inherits failures the model alone never shows: stale documents, broken tool calls, drift in live traffic.
The toolbox has standardized quickly. RAG systems are graded on the RAGAS metric family - faithfulness (does the answer follow from retrieved context?), answer relevancy, and context precision/recall - with the documented blind spot that these metrics score fidelity to whatever was retrieved, "not the reliability of your corpus." Agents add trajectory metrics: tool-call ordering, execution success rate, error recovery. Monitoring layers score every live response for faithfulness and hallucination, watching semantic drift (meaning shifts) separately from behavioral drift (decision changes). And online A/B testing closes the loop - with the practitioner warning that experiments "without power analysis is theater," and the deeper critique that engagement-optimized experiments may reward sycophancy over truth.
| Eval layer | Core question | Flagship example | Documented failure mode | Key mitigation |
|---|---|---|---|---|
| Benchmarks | What does the model know? | SWE-bench Verified: 60% to near 100% in one year (Stanford AI Index 2026) | Saturation + training-data contamination | Fresh-question sets (LiveBench); private held-out tests |
| Human review | Which output do people prefer? | Chatbot Arena blind pairwise votes, Bradley-Terry ratings | Vote noise, style bias, rigged votes shifting rankings (ICML 2025) | Anonymity, style control, confidence intervals |
| LLM-as-judge | Can AI grade AI at scale? | MT-Bench: ~80-85% agreement with human experts | Position, verbosity and self-preference bias | Swap-order averaging, rubrics, cross-family judges |
| Safety evals / red teaming | What could it do if misused? | Anthropic RSP (ASL-3, May 2025); OpenAI Preparedness; DeepMind CCLs | Self-authored thresholds; frameworks scored 10% on specificity | Third-party evals (METR, UK AISI); bug bounties |
| Regulation | What must be proven, by law? | EU AI Act Art. 55: standardized-protocol model evaluation (Aug 2025) | Compute thresholds are leaky proxies; enforcement still maturing | Documentation duties; incident reporting |
| Production monitoring | Is it still working now? | RAGAS faithfulness + context metrics; drift detection | "High benchmark scores routinely mislead teams" (arXiv 2026) | Custom eval sets from real traffic; CI gates; sampled human audit |
Adoption has outrun discipline: McKinsey's August 2026 survey found 40% of $1B+ organizations scaling AI agents, yet no rigorous survey exists on how many evaluate those systems properly - the honest answer is that this stat does not exist, and the gap is itself the finding.

Can a model ace benchmarks and still fail in production?
Yes, and the failure is structural rather than exceptional. Benchmarks measure a model on static, cleaned, pre-written tasks; production measures a system on messy live traffic, proprietary data and tool integrations the benchmark never saw. Contamination can inflate exam scores, benchmarks cannot see retrieval quality or tool reliability, and drift begins the moment deployment changes the input distribution. A February 2026 arXiv analysis concludes that high benchmark scores routinely mislead teams - which is why production stacks add custom eval sets from real traffic, regression tests in CI/CD, LLM judges with sampled human review, and drift monitoring on every live response.
The strongest counterargument, answered
The serious critique says the entire evaluation stack is closer to theater than measurement: benchmarks Goodhart themselves into uselessness (saturation plus contamination), human leaderboards have been shown gameable and structurally biased toward big labs ("The Leaderboard Illusion"), AI judges carry unremovable biases, safety frameworks are self-graded with thresholds set by the regulated party - DeepMind's framework scored 10% on specificity in one independent assessment - and production "evaluations" are mostly vendor dashboards selling reassurance. On this view, none of the five layers reliably tells you whether an AI system is actually good or safe.
The response grants the premise and rejects the conclusion. Every layer's failure mode is documented precisely because researchers inside the field measure the measurement - and those findings have already changed practice: style control on Arena, swap-order judging in CI pipelines, monthly fresh benchmarks, statutory evaluation duties in the EU, and the first eval-gated release cancellation in September 2026. More importantly, no better alternative exists at scale: the only substitutes for structured evaluation are vibes, vendor demos, or waiting for public failures. The stack is imperfect, partially gameable, and still the difference between testing a bridge before traffic and testing it after. The honest framing: AI evaluation is a gradually hardened approximation, not a solved problem - and knowing each layer's documented weaknesses is what lets you read the numbers responsibly.
What to actually do before shipping AI
For teams building with AI rather than grading frontier labs, the stack condenses into six moves, ordered by leverage.
- Build a custom eval set from your real traffic - 50 to 200 examples covering your actual use cases beat any public benchmark for your product, and become your regression baseline.
- Gate every prompt and model change in CI - rerun the eval set on each change; a score drop blocks the merge the same way a failing unit test does.
- Judge with a different model family, swapped order - self-preference and position bias are documented; use a judge from another provider, run both presentation orders, and average.
- Sample human review and measure agreement - audit a small percentage of judge verdicts with humans; if agreement drops below roughly 80%, fix the rubric before scaling automated judging.
- Monitor drift and faithfulness in production - score live responses for hallucination and track semantic versus behavioral drift separately; deployment changes the input distribution your evals never saw.
- Document everything now - if you serve EU users, GPAI-adjacent documentation duties are already law (Article 53 since August 2025); an eval paper trail is compliance armor as much as engineering hygiene.
FAQs
?How are AI systems evaluated?
Through a five-layer stack: standardized benchmarks (MMLU, SWE-bench, GPQA, Humanity's Last Exam) with automatic scoring; human preference evaluation (blind pairwise votes like Chatbot Arena, expert annotation); LLM-as-judge automated grading at scale; safety evaluations and red teaming before release; and production monitoring (drift, faithfulness, A/B tests) after deployment. Each layer has documented weaknesses - saturation, vote noise, judge bias, voluntary enforcement, benchmark-to-production gaps - so credible evaluation stacks all five rather than trusting one.
?What are AI benchmarks and why do they saturate?
Benchmarks are fixed question sets with automatic scoring run identically across models. They saturate because models improve faster than tests are rebuilt: the Stanford AI Index 2026 reports SWE-bench Verified performance rose from 60% to near 100% in a single year, MMLU passed 90% among frontier models, and on GPQA Diamond the top-model spread is smaller than the benchmark's own measurement noise. Contamination - test questions leaking into training data - accelerates saturation, which is why fresh-question benchmarks like LiveBench release monthly and still only reduce leakage rather than eliminate it.
?Is LLM-as-a-judge reliable?
Within limits. The MT-Bench research measured GPT-4 judges agreeing with human experts about 80% overall and 85% on non-tie comparisons, and a June 2026 study spanning 21 judges from nine providers is the largest validation to date. But three biases are documented: position bias (verdicts flip with answer order), verbosity bias (+0.24 to +0.44 effects measured length-aware), and self-preference (judges favor their own family's outputs even when humans rate them equal). Swapped-order averaging, explicit rubrics and cross-family judging make it dependable; raw scores without those controls are systematically distorted.
?How does Chatbot Arena ranking work?
Users type their own prompts, two anonymous models answer side by side, and the user votes for the better one. Millions of blind votes feed a Bradley-Terry model that produces the leaderboard's Elo-style ratings, with a style-control adjustment (added August 2024) discounting length and formatting bias. Documented criticisms include the "Leaderboard Illusion" paper's allegations of private multi-variant testing and selective disclosure favoring large labs, and an ICML 2025 study showing a few hundred rigged votes can shift rankings - so treat it as one signal, not ground truth.
?What is red teaming for AI models?
Structured adversarial testing where internal teams, contracted external experts, or paid bug-bounty hunters deliberately probe a model for dangerous failures - universal jailbreaks, harmful instructions, data leakage, rogue agent behavior. Anthropic pays up to $15,000 per novel universal jailbreak through HackerOne. Findings feed each lab's safety framework (OpenAI's Preparedness Framework, Anthropic's RSP with ASL thresholds, DeepMind's Frontier Safety Framework), and in September 2026 OpenAI withheld its GPT-6.1 Astra model after safety testing found unmet standards - the clearest case so far of evals gating a release.
?Are AI evaluations required by law?
In the EU, yes for general-purpose AI: since August 2, 2025, Article 53 of the AI Act requires all GPAI providers to maintain technical documentation, and Article 55 requires providers of systemic-risk models (presumed above 10^25 FLOP of training compute) to perform model evaluation under standardized protocols, run adversarial testing, and report serious incidents. The July 2026 Digital Omnibus deferred the separate high-risk-system rules to December 2027 but left these GPAI duties in force. In the US, the NIST AI Risk Management Framework remains voluntary guidance.
?How do you evaluate a RAG system?
With the RAGAS-style metric family: faithfulness (answers grounded in retrieved context), answer relevancy, and context precision and recall - scoring retrieval and generation as separate stages. Agent systems add trajectory metrics: tool-call ordering, execution success rate, and error recovery. The documented blind spot: these metrics verify fidelity to whatever was retrieved, not whether the corpus itself is correct, so a system can score perfectly while faithfully repeating a wrong source. Production adds drift monitoring and sampled human review on live traffic.
?Can a model ace every benchmark and still fail in production?
Yes - structurally, not exceptionally. Contamination can inflate exam scores; benchmarks never test your retrieval quality, tool integrations, or live traffic distribution; and a February 2026 arXiv analysis concludes high benchmark scores "routinely mislead teams" because evaluation pipelines carry their own silent failure modes. That is why production stacks add custom eval sets built from real traffic, CI regression gates, LLM judges with sampled human audit, and separate monitoring for semantic and behavioral drift after deployment.
The Bottom Line
How AI systems are evaluated in 2026 is a story of five imperfect layers compensating for each other's documented failures: benchmarks that saturate and leak, human votes that are noisy and gameable, AI judges that are cheap and biased, safety gates that are real but self-graded, and production monitoring that starts exactly where the exam ends. The field's honest position is that no single number - not one benchmark score, not an Arena rank, not a safety certification - tells you whether an AI system works. What separates a trustworthy evaluation from a marketing slide is whether the people behind it know which flaws to discount. The stack works when you use it the way this article does: layered, cross-checked, and always read with each layer's failure modes in mind.
Sources verified October 5, 2026. This article explains technical and regulatory developments for general audiences; it is not legal advice.
Sources
- hai.stanford.edu - Stanford HAI, The 2026 AI Index Report (SWE-bench saturation, HLE gains, benchmarks-are-saturating finding)
- arxiv.org - Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (NeurIPS 2023)
- arxiv.org - Wang et al., "Large Language Models are not Fair Evaluators" (position bias)
- arxiv.org - Panickssery et al., "LLM Evaluators Recognize and Favor Their Own Generations" (self-preference)
- arxiv.org - "The Leaderboard Illusion" (Arena private testing and disclosure critique)
- arxiv.org - Liu et al., "G-Eval: NLG Evaluation using GPT-4" (EMNLP 2023)
- arxiv.org - Ailem et al., distributional-robustness critique of benchmarks
- arxiv.org - OpenAI, "External Red Teaming for AI Models" (March 2025)
- arxiv.org - "Towards More Standardized AI Evaluation" (silent failure modes, Feb 2026)
- anthropic.com - Model safety bug bounty (up to $15,000, HackerOne)
- anthropic.com - Claude Opus 4 and first ASL-3 activation (May 2025)
- deepmind.google - Frontier Safety Framework (Critical Capability Levels)
- deploymentsafety.openai.com - OpenAI Preparedness Framework and GPT-5 system card (US CAISI + UK AISI evals)
- metr.org - METR third-party evaluations; "Common Elements of Frontier AI Safety Policies" (Dec 2025)
- artificialintelligenceact.eu - EU AI Act Article 55 (standardized-protocol model evaluation, systemic risk)
- artificialintelligenceact.eu - EU AI Act Article 51 (10^25 FLOP systemic-risk presumption)
- nist.gov - AI Risk Management Framework + Generative AI Profile
- mckinsey.com - The State of AI: Global Survey (Aug 2026): 40% of $1B+ orgs scaling agents
- pwc.com - PwC AI Agent Survey (May 2025): 79% executive adoption
- siliconangle.com - LiveBench launch (monthly fresh questions, Jun 2024)
- aclanthology.org - Chiang & Lee, human evaluation reproducibility (ACL 2023)
- paloaltonetworks.com - Red teaming definition
- docs.nvidia.com - RAGAS metrics reference (faithfulness, relevancy, context precision/recall)
- cset.georgetown.edu - AI red-teaming threat models and tools analysis


