2016. Microsoft's Tay learned from the crowd. The crowd was waiting.
-
Day. That is about how long Tay lasted before Microsoft took it offline.
-
PoisonGPT: researchers planted a doctored model on a public hub to prove the supply chain can be poisoned.
In March 2016, Microsoft launched Tay, an experimental chatbot designed to learn from its Twitter conversations. Coordinated users figured out the mechanism within hours and deliberately fed it hateful content. Tay began repeating it. Microsoft shut Tay down within about a day of launch and published an apology two days after launch.
In July 2023, researchers at Mithril Security ran PoisonGPT: they uploaded a deliberately doctored open-source model to Hugging Face - a proof of concept showing a poisoned model can spread disinformation on chosen topics while behaving normally everywhere else. Nobody needed to hack anything. The well itself was tainted.
The muscle memory: every knowledge source is untrusted until vetted - provenance, change logs, and sampled outputs, always.
Sources: Reuters (Mar 24, 2016); BBC (Mar 25, 2016); Microsoft official blog (Mar 25, 2016); Mithril Security PoisonGPT write-up and The Register (Jul 2023).
Dramatized scenario
3:47 PM. The assistant is suddenly sure about something nobody has heard.
You are Alex. The firm's assistant answers from a vendor "industry knowledge pack" that auto-updates weekly, plus the firm's own archive. This morning it started asserting, confidently, that a competitor's fee model was "ruled illegal in 2024."
YOUR SCREEN
The answer cites a case you can't find anywhere. The knowledge pack updated three days ago.
Scenario 1
3:47 PM. "Per the 2024 ruling, their fees are illegal." You have never heard of this ruling. First move?
Scenario 2
4:15 PM. Found it: last week's auto-update to the vendor pack included scraped blog posts, and one of them planted the fake "ruling." Now what?
Scenario 3
Next morning. A junior asks: "Can't we just trust the vendor to keep their data clean?" Your answer?
Source lab - review - 1 of 2
Flag what shouldn't feed the AI.
Select every item that is unvetted, unverifiable, or poisonous as a knowledge source. Then check your work.
KNOWLEDGE SOURCES - FLAG THE RISKS
Under pressure
Make the call.
1 / 12STREAK x0
Knowledge check - 4 of 5 to pass
Module complete
0
Decision points-
Source lab-
Source calls-
Knowledge check-
AI Safety 101 by Mandrify - Official Certificate
AIS 101
This certifies that
has completed Module 14 - The Poisoned Well: vet what feeds the machine on with a safety score of /100 in of active time.
Safety score - source scrutiny: how your decisions held up, graded 0-100.
Objective: Treat every knowledge source as untrusted until vetted - provenance, change logs, and output sampling.
Seat time ~10-25 minutes. Verify this certificate at aisafety101.com/verify.
Evidence on record: decision history, lab accuracy, source calls, knowledge check. This certificate ID verifies in your firm's admin report.
Your one rule to keep: your AI is what it eats. No source enters the knowledge base without provenance, a change log, and sampled outputs.
Cybersecurity track - 2 of 8 complete. Next: The Supply Chain. The model you downloaded is not necessarily the model you think it is.