8 total
AI Benchmark A standardized test that runs the same inputs through one or more AI …
Added 1 Jul · Upd 1 Jul ·6 min
AI Evaluation The whole practice of judging whether an AI system is fit for use, …
Added 29 Jun · Upd 29 Jun ·4 min
AI Safety What AI safety is, the categories of harm it addresses, and the … Glossary
Added 28 Mar · Upd 30 May ·3 min
How AI Models Are Evaluated: The Hidden Lifecycle The invisible pipeline behind every chat box: training, alignment, red … Guides
Added 1 Jul · Upd 1 Jul ·9 min
Meta secretly benchmarked rival chatbots by posing as teens A WIRED investigation found Meta ran covert safety testing against … News
Added 1 Jul · Upd 1 Jul ·4 min
Model Evaluation Model evaluation tests an AI model in isolation, measuring its raw …
Added 29 Jun · Upd 29 Jun ·4 min
Red Teaming What red teaming is in AI, how adversarial testing discovers … Glossary
Added 28 Mar · Upd 30 May ·4 min
Red Teaming and Adversarial Testing for AI Systems How to plan and execute red team exercises that systematically probe AI … Guides
Added 28 Mar · Upd 30 May ·3 min