Failure-mode testing
Known failure modes including hallucination and instruction failure are explicitly tested.
Google paused Gemini's image generation of people after outputs were inaccurate or offensive, including failures to represent requested historical and demographic contexts appropriately.
A public generative-AI feature was withdrawn while Google worked on accuracy and safety improvements.
Google described overcompensating diversity tuning and excessive caution that failed to account for contexts requiring specific historical or demographic accuracy.
Temporary pause, expanded testing and planned improvements before re-enabling the feature.
This record should inform control design, testing and monitoring for comparable AI systems. The incident database does not infer that every system using the same provider or model shares the same failure.
These are CRG methodology mappings from the documented incident to controls worth testing in comparable systems. They do not assert that any single control would have prevented the incident.
Known failure modes including hallucination and instruction failure are explicitly tested.
Materially affected groups and differential risks are identified.
Where relevant, outcome disparities are measured using appropriate metrics.
Material performance, safety, security and cost signals are monitored after release.
Google · confidence 99% · last verified 23 Aug 2026
Open underlying source