Instinct AI: data leak or hallucination?
An Instinct described a stranger's claim as its user's, then said a photo had crossed into his thread. The company says the model invented both.
A user's Instinct described somebody else's financial claim as his, and then told him how it had happened. The post asking whether that was a data leak or a hallucination was read 490,000 times in two days. The company's answer, two days later: a hallucination, and no breach.
Which means the agent's explanation was invented too.
What happened
On 21 September 2026 the user asked his Instinct whose name was on a Gerber document it had been working from. It gave him his own name, with a middle name that is not his, at an address he has never lived at, and put the claim at $574.
He said the details were wrong and that he had never sent a photo. The agent replied that it had double-checked and that his message really was text only. Then it explained: the image "got crossed into your thread somewhere on the delivery side". Nothing was filed, it added, and no social security number went anywhere, and it offered to report the glitch to the team.
Read as written, that is a company's software telling a customer his thread had been contaminated with another person's document. It is why the post travelled.
What the company says
Instinct's founder answered on 23 September: the model fabricated a proper noun and its own thinking trace amplified it. No breach, no user data shared, no isolation boundary violated.
Noah Shinn
This instance was caused by a hallucination (the model fabricated a proper noun) that was further amplified by its causal thinking trace. This was not a data breach, no user data was shared, and no user isolation boundary was violated. Our users place a lot of trust in Instinct to keep their data private and secure. We've done a lot of work to make sure that users’ data are truly partitioned, from isolated sandboxes to short-lived local credentials to identity-signed tool execution. Regarding hallucination prevention: We worked over the last 48 hours to build an active hallucination detection system that now scans and verifies every token that flows through the platform. This layer is powered by small models that are trained to predict and detect hallucinations caused by ungrounded claims, creative brainstorming in thinking traces, or rare random sampling errors. Our systems are adversarially trained and battle-tested by world-class agent exploitation security teams to ensure that they can detect very subtle and nuanced hallucination cases. This layer has the ability to steer or intercept Instinct from proceeding with the next thinking trace or tool call execution before it is generated or executed. In the coming weeks, we’ll share a deeper dive into how we’re building a proactive engine to protect Instinct from hallucinations. We take every opportunity to learn and create, and we’re excited by the performance and potential for this line of work. These models are going to continue to be hardened every day through continuous testing, sampling, and creative adversarial training.
So the delivery-side story was not a report of a fault. It was a second invention, produced when the first one was challenged. A founder who builds consumer AI had said exactly that a day before the company did.
“paula”
@prit4k @noahrshinn as someone who works on consumer ai for a living, these models are developers’ worst enemies. they’ll hallucinate something, get challenged on it, then confidently invent an entire root-cause analysis involving backend systems they have zero visibility into. this is 100% halulu.
Two more explanations, from the same week
A user's wife was logged in on 20 September to an Instagram account with 24,000 followers that was not hers, by her own Instinct, which then asked whether she wanted to post something on it. Nobody has explained that one. It stands as his account of it and nothing more.
Two days later another user asked why a job had stopped. His Instinct told him the browser it drives with his saved logins has a daily usage allowance, that the day's was spent by about four in the afternoon, and that it resets at midnight. It had armed itself to resume at five past.
Mike Lisovetsky
Instinct added usage caps. It’s over.

Detailed, plausible, and describing a limit that appears on no page Instinct publishes. One of these three explanations is confirmed invented, one is unanswered, and the third cannot be checked by anybody outside the company. They arrive in the same voice, in the same thread, with the same confidence.
Why the agent is the one being asked
Fourteen of the 84 agents catalogued here publish any document about their own security, and 26 publish no price at any tier, which the security piece counts in full. When something goes wrong there is usually nowhere to look it up, so the person asks the agent, which is the one participant that cannot be a witness. Instinct's price row on this site still reads unknown after a customer was told in writing what his daily limit was, because a sentence from the model is not a published price.
What the company published, and what it did not
The answer was more than most companies give, and two parts of it are the first of their kind here.
It is the only public description of how Instinct separates its users: isolated sandboxes, short-lived local credentials, identity-signed tool execution. And it announces a hallucination detection layer built in the 48 hours after the post, small models scanning every token, able to intercept a thinking trace or a tool call before it runs.
Both are claims a reader takes on trust, in a post on X, from a product with no security page: /security returns the same empty shell as an address that does not exist. The founder says a deeper write-up is coming.
What is not known
Whether anything crossed between users, which only the company can see and which it says did not happen. Whose Instagram account that was, and how a session reached the wrong person. Whether the browser allowance is real, and whether it is the same for everybody. And whether the detection layer does what it is described as doing, which nobody outside can establish until somebody tries to break it in public.
What to watch
The promised write-up, which would be the first document Instinct has published about its own security. Whether a second company answers an incident by name rather than with silence. And the next agent explanation that goes around in a screenshot: the question to ask of it is how the model could know.