“There is nothing either good or bad, but thinking makes it so.”
– William Shakespeare, Hamlet (Act 2, Scene 2).
This may sound a bit crazy, but AI can be trustworthy. (No, AI did not write this :>)
What if the risk of AI is not inherent in the technology? What if instead, it is an inheritance, an amplification of the cracks and underlying faults in our thinking? Generative AI output is the product of the financial, societal, legal, and practical frameworks that we use to build, train, deploy and use AI today. Recent studies by OpenAI “Why Language Models Hallucinate”and“Predictable Compression Failures: Why Language Models Actually Hallucinate,” show that hallucinations, the most egregious trust violations, are a byproduct of the way we train models and process data.
Of course hallucinations are not the only risk AI adds to our work, institutions and lives. Beyond accuracy and confidence, we see many more risks related to security, bias, privacy and the most mysterious frontier; explainability.
What if the models themselves are not even at the heart of these AI safety risks? To explore this question, let’slook at three factors that play a role in AI development; human behavior, user behavior, and model behavior.
What, Me Singularity?
First, we need to envision a timeline where things don’t go full Skynet. In this universe, AI doesn’t hallucinate, help hackers exploit vulnerabilities, or spiral into cascading complexity.
The good news is that we already have a number of institutions dedicated to maintaining the good timeline.
MITRE is the architect of coherance. The AI Assurance Guide strives to confirm that AI systems behave consistently across time, context, and scale.
NIST is not going to tell you that time travel is not possible. However, if you decide to take the risk, NIST AI Risk Management Framework will help you measure, monitor and admit when you accidently showed up in spider web, wearing a fly’s head.
Stanford’s HAI, (Institute for Human-Centered Artificial Intelligence), gives us a map of the known universe. HAI tracks what is actually happening across companies, industries and societies.
MIT provides the warp core of understanding. MIT’s Responsible AI initiative establishes that AI’s power is meaningless without interpretability. Machines thinking fast without human reflection is unstable and unsustainable.
And finally, OWASP knows that the bad bots are already out there. The Top 10 for LLMs is one of the tools the resistance uses to take out the terminators and prompt manipulators searching for your security vulnerabilities.
But I thought we were talking about Hamlet, not John Conner? Well, both, and neither.
“We can be blind to the obvious, and we are also blind to our blindness.”
– Daniel Kahneman, Thinking, Fast and Slow
Human Behavior
That Doggone Availability Heuristic
Nobel Prize Winner DanielKahneman’s book title, ‘Thinking Fast and Thinking Slow’, indicates two systems humans use in cognition. System one is the instant, emotional, intuitive thinking; and system two is slower, reasoning, deliberate thinking. Throughout Kahneman’s extensive research he concludes that we rarely notice our own cognitive mistakes. Even worse, we do not realize that we don’t notice! Because our intuitions feels right, we trust them.
Daniel Kahneman on the Lex Fridman Podcast.
Heuristics are tools or mental shortcuts or rules of thumb that allow people to make quick decisions and judgments with minimal mental effort.
For example, if someone says ‘What did you eat for lunch?’,
and you see the word fragment
S0_P
Your brain will fill in the missing letter differently than if you see or were just talking about cleaning the bathroom.
SO_P
Your system one decides for you.

The same way that your brain assigns autonomic systems to continue functioning while it gives up on more advanced calculation can be applied to foundation models.
Transformers and Prime
In psychology, ‘priming’ happens when exposure to one stimuli directly influences your response to another stimuli. Human actions and emotions can be primed by events we are not even aware of.
Priming events are not restricted to words. Environmental cues can change our motivations, our mood, or behaviors, without us even noticing.
Subtle reminders of money unconsciously shift people into a self-sufficient, task-focused mode that increases performance but decreases social warmth, generosity, and willingness to depend on others.
In her research on money priming, Kathleen Vohs discovered that even subtle reminders of money unconsciously shift people into a self-sufficient, task-focused mode. This increases their performance, but decreases social warmth, generosity, and willingness to depend on others.
It is possible that priming may be true for gen AI systems. Research on LLM shutdown resistance and Anthropic’s 16 Model Self Preservation Study indicate that these threatening behaviors may be a result of how the model is trained.
Don’t Be So Optimistic
Flash back to the University College London. Its 2011 and Tali Sharot has just published her study “How unrealistic optimism is maintained in the face of reality.”
In this study a score of students were asked what they believed their likelihood of getting cancer, Alzheimers and other traumatic life events was. Participants were then shown the actual statistical probability that they would suffer from each of those afflictions. Oh, and Dr. Shalot conducted this test while they were in an fMRI machine.
You may not be surprised to learn that the study showed that they (we) are optimistically biased. The participants left inferior frontal gyrus (the part of our brain that overrides system one in Kahneman’s research) showed greater updating activity when the results were better than they expected, and reduced updating when news was worse than expected.
It is starting to look like transformers hallucinate LESS than humans! :>
User Behavior

Hackers, Attack!
How many hack attempts does Amazon endure every single day? 10,000? 100,000?
How about 750 Million!
In 2024 Amazon reported that the number of daily attacks went from 100M to 750M. In large part this 7.5x increase is due to autonomous attacks made possible by the proliferation of AI tools.
Users are a vulnerability and will continue to expand the threats across a wider and wider interaction surface.

Enterprise Buyers and the New ROI
SAS and technology security historically has been a linear process based on code deployment. Endpoints are contained and vulnerabilities have specific test criteria QA’d during development. Cybersecurity provided value by identifying when a system was being hacked or broken, creating a patch for it, and verifying the results. This provides accounting and measurement of thwarted attacks. It’s assuring.
The buying muscle that executives have strengthened under this standard procedure has a few fundamental flaws when it comes to AI systems. First, if an AI system breaks in the world, then reputational damage and breach of trust has already occurred, regardless of how well you fix the underlying bots. Secondly, the fix for one type of AI vulnerability does not mean that you won’t have a different harm occur. While this is true in software security as well, the risk has multiplied by the ability of regular users to inadvertently break the system.
CISO and C-Suite teams have to rethink how ROI works with AI. Take responsibility for the output of AI deployment, regardless of the number of AI models and systems included in the solution. Preventing financial, legal, physical, psychological and societal harm requires new levels of transparency and data sharing.

Regular Users and Outdated Fine Print
Legal system based on punishment and reimbursing victims after a break occurs is not a prudent or effective strategy when deploying AI. While it will be necessary to hold foundation model makers, wrapper companies, and AI systems developers accountable, accountability needs to focus on prevention much more so than it does consequences and penalties.
In the financial sector, there are multiple cases were penalties even in the hundreds of millions of dollars are an “acceptable loss“ when compared to the value or saving created by an AI system. Examples include AI system deployments in trading algorithms, loan application system, or other financial transaction software. In order to have truly safe AI or at least certifiably safe AI we need to rethink what the public trust mechanism, reputational impact and public contract means in an age where chat bots, AI agents, LLMs, NLP’s and other AI systems are deployed without transparency or regulation.
Model Behavior

Model, Heal Thyself
Which brings us to the model’s role in being trustworthy. We’ve seen major progress both in explainability and hallucination reduction.
We have also seen models like Grok, OpenAI Sora, Gemini and others continue to fail in terms of bias and supply information that can be harmful to regular users. Additionally their are dozens of models accused of Privacy violations with active litigation against Anthropic, Otter.ai, Clearview.ai and more.
Frameworks such as OWASP, NIST, IBM Ai Fairness 360 and model-level guardrails like Amazon bedrock, Anthropic’s Constitutional AI, and Gemini Safety Filters can be effective when properly deployed in consumer facing models. The next step in model safety will e to mitigate responses through predictions as we see the principles of these frameworks eroding in real time.
Isolated testing practices will never be able to keep up with the occurrence of new conditions across widening spectrums of interactions between AI agents and systems. This will require new practices and organization. I’m testing methodologies in the emergence of new tools which are inference and compute optimized to work alongside the performance of models.

Alas LLM, I Knew You Well
Recent developments in explainability are encouraging. The University of Oxford (UK) and the Indian Institute of Information Technology integrated Explainable AI, (XAI) with the OD-DRUNN threat classifier to create proactive cybersecurity defense. This approach improves explainability by tying each alert to the exact flows and users involved. This gives the teams immediate insights into why did this happen and why now.
To learn more on what is to be or not to be (It had to be done) for human understanding and explainability, check out this Paperguide.
Conclusion
AI disasters do not have to be inevitable. In order to create a future where AI and humans work harmoniously, we need to look beyond what algorithms and models are computing. We need to value preparation based on open shared information and model disclosure instead of a punitive system that promotes a lack of clarity. Strategic preparedness is more valuable than failed attempts to control outcomes. As Hamlet concludes, ‘ The readiness is all.’

