Impact assessment
Material impacts on users, workers, customers or the public are assessed.
Investigative testing found New York City's MyCity chatbot giving incorrect answers about housing, employment and business rules, including advice inconsistent with applicable law.
The incident raised concerns about using generative AI as an authoritative public-service information interface without sufficiently reliable grounding and escalation.
Generative responses were presented in an official government context while factual/legal verification and confidence gating were insufficient for some questions.
The city added stronger warnings and continued work on improving the beta service.
This record should inform control design, testing and monitoring for comparable AI systems. The incident database does not infer that every system using the same provider or model shares the same failure.
These are CRG methodology mappings from the documented incident to controls worth testing in comparable systems. They do not assert that any single control would have prevented the incident.
Material impacts on users, workers, customers or the public are assessed.
Known failure modes including hallucination and instruction failure are explicitly tested.
Research/advice systems preserve source provenance and distinguish evidence classes.
Known limitations and out-of-scope uses are communicated to operators and users.
The Markup / THE CITY · confidence 95% · last verified 23 Aug 2026
Open underlying source