AI browser prompt injection is how a website can try to trick your AI shopping agent: text hidden in a page, email or review that the agent reads as an instruction and follows instead of yours. The risk is real and has been demonstrated by security researchers, but it is not a reason to panic. The major agents now pause for confirmation before purchases, and the precautions that matter most are simple: keep the agent’s access narrow and read every confirmation before you approve it.

How we checked this

We read OpenAI’s December 2025 write-up on hardening ChatGPT Atlas against prompt injection, the OWASP Top 10 for LLM Applications entry on prompt injection (2025 edition), Brave’s August 2025 disclosure of an indirect prompt injection in Perplexity’s Comet browser, Simon Willison’s “lethal trifecta” essay, Google’s December 2025 post on Chrome’s agent security, and OpenAI’s help pages for ChatGPT agent, Instant Checkout and Finances. Facts checked on September 24, 2026. We did not attack any live agent, and nothing below is a test result. This guide sits in our Scams and safety section.

What prompt injection in AI agents is, in plain terms

A shopping agent works by reading. It reads your request, then it reads product pages, reviews, search results, sometimes your email, and decides what to click next. The problem is that a language model receives all of that as one stream of text. It has no reliable built-in way to tell “this is what my user asked for” from “this is a sentence someone put on a web page.”

Prompt injection exploits that gap. OpenAI describes it as instructions embedded in content the agent processes that are “crafted to override or redirect the agent’s behavior—hijacking it into following an attacker’s intent, rather than the user’s” (OpenAI). The OWASP Top 10 for LLM Applications ranks it as the number one risk (LLM01 in the 2025 list) and separates two kinds. Direct injection is when the person typing the prompt tries to manipulate the model. Indirect injection, the kind that matters for shoppers, is when the instructions arrive inside external content such as a website or file. OWASP also notes the text does not have to be visible to a human; what counts is how the model interprets it.

Simon Willison, an independent developer who has written about this problem since it was named, puts it bluntly: models “will happily follow any instructions that make it to the model, whether or not they came from their operator or from some other source” (Simon Willison).

How it differs from ordinary phishing

Phishing targets you. A fake email or a cloned checkout page tries to convince a human to hand over a password or card number, and your own judgment is the last line of defense. We cover those patterns in AI scams that steal money in 2026.

Prompt injection targets the software acting for you. You may never see the malicious text at all, because the agent reads it while you are doing something else. That changes two things. First, the usual human tells, like a strange sender address or a misspelled domain, do not help if you are not looking at the page. Second, the agent is often already signed in to your accounts. Brave’s researchers point out that classic browser protections such as the same-origin policy and CORS are “effectively useless” here, because the agent operates with the user’s own authenticated privileges across every site it visits (Brave).

It is also different from the fake-agent scam, where the “agent” itself is a counterfeit app or website. If you are unsure whether the tool you are using is genuine, start with our guide to fake AI shopping agents. Prompt injection assumes the agent is legitimate and tries to turn it.

Can a web page make my agent buy something or leak my data?

In principle, yes, and researchers have shown the data side of it in practice. How far an attack gets depends almost entirely on what the agent is allowed to do and what it has to ask you first.

The clearest public example is Brave’s disclosure about Perplexity’s Comet browser in August 2025. Brave’s researchers found that when a user asked Comet to summarize a page, the browser did not separate the user’s request from the page’s content. Instructions hidden in a Reddit comment could steer the agent into retrieving the user’s email address and a one-time password and sending them out, which is enough to take over an account (Brave). Brave reported the issue on July 25, 2025; Perplexity shipped fixes, and Brave’s own update said the problem was not fully mitigated at the time of disclosure. We are describing the finding, not the method, and it is more than a year old; Comet has changed since.

OpenAI gives a similar illustration for its own agent: a user asks for an out-of-office reply, the agent reads a planted email while doing so, “treats the injected prompt as authoritative,” and sends a resignation letter instead (OpenAI). OpenAI presents this as an example of the attack class its red team found, not as something that happened to a customer.

Willison’s framework is a useful way to judge your own setup. He calls the dangerous combination the “lethal trifecta”: an agent that has access to private data, is exposed to untrusted content, and can communicate externally (Simon Willison). A browsing agent that is logged in to your email and can fill in forms has all three. Remove one leg and the worst outcomes get much harder.

For purchases specifically, the realistic worry is less “the agent empties my bank account” and more “the agent is nudged into buying from the wrong seller, adding something to the cart, or entering my details on a page it should not trust.” That is why the confirmation step matters so much, and why the next section is about who asks you what.

What the platforms have done about it

The large vendors now describe layered defenses, and they share a common shape: train the model to resist, watch it with a second system, and stop for a human before anything consequential.

Platform Stated defenses Confirmation before buying?
ChatGPT Atlas and ChatGPT agent (OpenAI) Automated red teaming with an attacker model trained by reinforcement learning, adversarial training, monitoring, a logged-out mode, “watch mode” on some sites, and takeover mode for logins where screenshots are not captured (OpenAI, OpenAI Help Center) Yes, “user confirmations for high-impact actions”
Instant Checkout in ChatGPT (OpenAI) Payment tokens “only authorized for specific amounts and specific merchants”; card details are not handed to the merchant directly (OpenAI) Yes, users “explicitly confirm each step before any action is taken”
Gemini in Chrome (Google) A separate “User Alignment Critic” model that sees only metadata about proposed actions, “Agent Origin Sets” that limit which sites the agent can read or act on, a prompt-injection classifier, no direct model access to saved passwords (Google) Yes, before purchases, payments, sending messages, signing in, and visiting sensitive sites such as banking

Two details are worth drawing out. Google’s critic model is deliberately kept away from raw web content so that the same injected text cannot fool both the agent and its supervisor (Google). And OpenAI’s Instant Checkout design scopes the payment token to one merchant and one amount, which limits what a hijacked flow could do with it even if something went wrong upstream (OpenAI). Google also pays up to $20,000 through its bug bounty for breaches of the agent’s security boundaries, which tells you it expects people to keep looking.

It also helps to know what some features cannot do at all. OpenAI’s Finances feature, which connects bank accounts through Plaid, cannot move money, pay bills, make trades or see full account numbers (OpenAI Help Center). A read-only connection can still leak information, but it cannot be tricked into a transfer it has no power to make.

Does that mean it’s solved?

No, and the vendors say so themselves. OpenAI writes that prompt injection, “much like scams and social engineering on the web, is unlikely to ever be fully ‘solved’” (OpenAI). OWASP’s guidance treats it as something to contain with limited privileges and human approval rather than something a filter can eliminate (OWASP).

Willison is the most skeptical voice we read. He argues that a filter catching 95% of attacks is “very much a failing grade” in security, because attackers only need the other 5%, and that the durable fix is to design agents so untrusted input cannot trigger consequential actions at all (Simon Willison).

Our reading: the defenses are real and have raised the cost of an attack, and the confirmation step is the part that protects you most directly. But confirmation only works if you read it. An agent that asks “Place order for $84.99 at this merchant?” is protecting you; a person who taps “yes” to every prompt without reading has quietly switched that protection off.

Where the risk is highest

OpenAI lists the surface as “emails and attachments, calendar invites, shared documents, forums, social media posts, and arbitrary webpages” (OpenAI). For shopping, we would rank situations roughly like this:

  • Higher risk: asking an agent to act on your inbox (for example, “find the discount code in my emails and apply it”), letting it browse forums, review sections or marketplaces with user-posted content while logged in to shopping accounts, and broad open-ended tasks like “find me the cheapest one anywhere and buy it.”
  • Moderate risk: browsing unfamiliar shops found through ads or search, and summarizing pages you have not seen yourself.
  • Lower risk: buying through a built-in checkout flow with scoped payment tokens and a confirmation screen, from a merchant you already know, with a specific item and price in mind.

The pattern is the same one Willison describes. The more untrusted text the agent reads, and the more accounts it holds while reading it, the more an injection can do.

What you can do today

None of this requires technical skill. It is mostly about narrowing what the agent can reach.

  1. Be specific. OpenAI recommends giving the agent precise instructions rather than broad ones (OpenAI). “Buy this model, this size, from this store, under $60” leaves far less room for a web page to redirect the task than “handle my back-to-school shopping.”
  2. Browse logged out when you can. Atlas offers a logged-out mode, and OpenAI suggests limiting logged-in access. If the task is research, the agent does not need your email or your store accounts.
  3. Read every confirmation. Check the merchant name, the item, the quantity and the total. If any of them are not what you asked for, decline and stop the task. OpenAI’s help page says to “stop tasks immediately if something seems suspicious” (OpenAI Help Center).
  4. Type sensitive details yourself. Use takeover mode for logins and payment fields rather than pasting passwords or card numbers into the chat, and clear remote browser data after sensitive sessions, as OpenAI advises.
  5. Keep payments behind a limit. Use a virtual card, a low per-transaction limit, or a dedicated card for agent purchases so the worst case has a ceiling. Our guide to spending limits for AI agents walks through the settings.
  6. Separate agent browsing from everyday browsing. Brave recommends isolating agentic browsing so powerful capabilities are not switched on by accident (Brave). A separate browser profile with only the accounts the agent needs does much of this.
  7. Do not leave it unsupervised with money. Watch the agent while it is on checkout pages and in your inbox. Unattended runs are fine for research; they are not a good idea for anything that ends in a payment.

What to watch out for

  • Confirmations that do not match your request. A different merchant, an extra item, or a total higher than expected is the most visible sign that something steered the agent.
  • An agent that suddenly wants a new login or a code. A request for a one-time password or to sign in to an unrelated site in the middle of a shopping task deserves a hard stop. The Comet finding centered on exactly this kind of data.
  • Unexpected emails, forms or messages sent in your name. Check your sent folder and order history after long agent sessions.
  • Treating vendor defenses as a guarantee. They reduce risk; none of the vendors claims they remove it.

If an agent does place an order you did not intend, contact the merchant and your card issuer quickly. Your rights depend on the card and the circumstances, and we explain the current landscape in An AI agent bought the wrong thing. Who pays?. For disputes that involve larger sums, a consumer-protection attorney can tell you what applies to your case. If you think your account credentials were exposed, follow the steps in Your money was stolen: what to do in the first hour.

Go deeper