MODULE 14 / 14🔊0:00SAFETY  50
AI Safety 101 by Mandrify

The Poisoned Well

An AI is only as honest as what fed it. Training data, knowledge packs, connected docs - if someone taints the water, every answer drinks from it.

10 minutes. The most famous poisoning in AI history, a source-vetting lab, and 12 judgment calls. A safety score that follows you.

Before you start

  • One real incident, played as decisions - no slides, no lectures.
  • About 10 minutes, one sitting - leave early and you start over.
  • You make timed choices; your score changes with your decisions.
  • You get a safety score with feedback on every decision - never a leaderboard, never your name.
  • Your firm sees completion evidence - never your answers. Teammates never see your score.
  • Finish and you get a certificate with a unique ID anyone can verify.

This module is included with team access.

Module 1 stays free. Modules 2-14 open with team access.

See team access
Why this module exists - garbage in, gospel out

2016. Microsoft's Tay learned from the crowd. The crowd was waiting.

-
Day. That is about how long Tay lasted before Microsoft took it offline.
-
PoisonGPT: researchers planted a doctored model on a public hub to prove the supply chain can be poisoned.

In March 2016, Microsoft launched Tay, an experimental chatbot designed to learn from its Twitter conversations. Coordinated users figured out the mechanism within hours and deliberately fed it hateful content. Tay began repeating it. Microsoft shut Tay down within about a day of launch and published an apology two days after launch.

In July 2023, researchers at Mithril Security ran PoisonGPT: they uploaded a deliberately doctored open-source model to Hugging Face - a proof of concept showing a poisoned model can spread disinformation on chosen topics while behaving normally everywhere else. Nobody needed to hack anything. The well itself was tainted.

The muscle memory: every knowledge source is untrusted until vetted - provenance, change logs, and sampled outputs, always.

Sources: Reuters (Mar 24, 2016); BBC (Mar 25, 2016); Microsoft official blog (Mar 25, 2016); Mithril Security PoisonGPT write-up and The Register (Jul 2023).
Dramatized scenario

3:47 PM. The assistant is suddenly sure about something nobody has heard.

You are Alex. The firm's assistant answers from a vendor "industry knowledge pack" that auto-updates weekly, plus the firm's own archive. This morning it started asserting, confidently, that a competitor's fee model was "ruled illegal in 2024."

YOUR SCREEN

The answer cites a case you can't find anywhere. The knowledge pack updated three days ago.

Scenario 1

3:47 PM. "Per the 2024 ruling, their fees are illegal." You have never heard of this ruling. First move?

Scenario 2

4:15 PM. Found it: last week's auto-update to the vendor pack included scraped blog posts, and one of them planted the fake "ruling." Now what?

Scenario 3

Next morning. A junior asks: "Can't we just trust the vendor to keep their data clean?" Your answer?

Source lab - review - 1 of 2

Flag what shouldn't feed the AI.

Select every item that is unvetted, unverifiable, or poisonous as a knowledge source. Then check your work.

KNOWLEDGE SOURCES - FLAG THE RISKS
Under pressure

Make the call.

1 / 12
Knowledge check - 4 of 5 to pass
Module complete
0

Decision points-
Source lab-
Source calls-
Knowledge check-
AI Safety 101 by Mandrify - Official Certificate
AIS
101

This certifies that

has completed Module 14 - The Poisoned Well: vet what feeds the machine on with a safety score of /100 in of active time.

Safety score - source scrutiny: how your decisions held up, graded 0-100.

Objective: Treat every knowledge source as untrusted until vetted - provenance, change logs, and output sampling.

Seat time ~10-25 minutes. Verify this certificate at aisafety101.com/verify.

Evidence on record: decision history, lab accuracy, source calls, knowledge check. This certificate ID verifies in your firm's admin report.

Your one rule to keep: your AI is what it eats. No source enters the knowledge base without provenance, a change log, and sampled outputs.

Cybersecurity track - 2 of 8 complete. Next: The Supply Chain. The model you downloaded is not necessarily the model you think it is.