LLMs (large language models) are like cats
If you've been on a webinar with me since 2014, you probably have seen my longest tenured co-worker: a domestic, short-haired rescue cat that my sons named Pancake when we adopted him that year. Although I can't say that he contributes much to my work, he works for a low salary and is often entertaining when he joins my Teams calls.
He's been variously described as “majestic” (his vet), “beautiful” (my daughters), “disdainful” (my sons), “entitled” (me), and fuzzy.
I was recently speaking at a customer dinner about the risks of AI and why Keepit is investing in building our AI Truth Cloud solution. I had just pointed out that one major problem with today's generative AI solutions is that you can give them instructions, and they may or may not follow them. The only guaranteed way to prevent an LLM from doing something that you don’t want it to is to prevent it from being able to take that specific action. Guardrails don't work reliably enough.
As I was standing there, microphone in hand, it dawned on me: LLMs are like cats.
How LLMs are similar to cats
You can train them in various ways and reward them for desirable behavior. And yet, no matter what instructions you give them, there is a non-zero chance at any point in time that they will do just exactly whatever they feel like doing instead of what you told them to.
As evidence, let me share a cat story. A couple of years after we got him, and for no good reason, Pancake decided that the tub in my bathroom made a great replacement for his litter box. No amount of aversion or positive reinforcement would get him to quit fouling the tub (good thing I prefer showers!).
The only thing that worked was making sure to keep the bathroom door firmly closed so he couldn't get to it.
Pancake when he's being good.
Why blocking access (not guardrails) is the only reliable way to prevent AI-driven data loss
I think the parallel to the way that we work with LLMs is pretty obvious. Every AI company has invested heavily in trying to improve the alignment of their various models. The only tactic that works reliably well is to block the model's ability to modify or delete data that you don't want it to touch. Blocking doesn't necessarily help with alignment problems related purely to output, such as trying to keep the LLM from telling people how to make methamphetamine or launder the profits thereof, but it can stop the commonplace occurrence that a misbehaving model or poorly chosen or malicious prompt will cause data loss.
The cat-like nature of LLM-based AI tools is a huge motivation behind our AI Truth Cloud. We recognize the tremendous benefit that adding AI to your organizational playbook can bring. It can help you do more of the right work, faster, and with better quality… as long as it's behaving. When it doesn't behave, the AI Truth Cloud can help you restore your production systems quickly to the correct state to keep your business going.
Where do we go from here?
We'll have a lot more to say about the AI Truth Cloud as we continue to refine and develop it. Meanwhile, you might check out our catalog of webinars, because if I'm speaking, there's an excellent chance that you'll get to see Pancake in all his native majesty.