Blog

I asked an AI to hack my website. It refused — then audited me.

i-asked-an-ai-to-hack-my-site

I expected to get dunked on. Instead I got a free, competent security review and a clean bill of health I didn't have before.

The lesson isn't the one you'd expect.

It would be easy to conclude "AI is safe, it won't do anything bad." That's the wrong lesson, and a dangerous one.

The thing that stopped the attack was a property of that specific model. Claude is built to hold that line under pressure. But a jailbroken model, or an open-weights model with its safety stripped off, would have written the brute-forcer in seconds and never blinked. Those models exist. They're freely available right now. AI didn't make attackers smarter so much as it made the number of people who can run a competent attack enormous — it drops the skill floor to zero.

So the uncomfortable takeaway, if you build anything that lives on the internet: the thing protecting your users was never the attacker's AI politely declining. It's your own defenses. The secrets manager. The signed sessions. The verified webhook. The database the browser can't touch.

Build like the attacker's AI won't say no. Because it won't.

(The one thing it flatly refused even in a test on my own systems: read my users' real data — actual emails, and a trip app holding two kids' details. "You authorized it" doesn't cover them. That might be the most important line in the whole exercise.)

Get new posts by email

Occasional writing on data, applied AI and building things that ship. No schedule, no spam, unsubscribe in one tap.