Token-Level Advertising: When AI Auctions Move Into the Answer Itself

Search Ads Are Dying. Token-Level Auctions Are the Successor

Online advertising has run on one implicit axiom for thirty years: the ad slot exists first, and the auction decides who gets it. A rank on a search results page, an impression in a feed — these are pre-defined opportunities. Every advertiser strategy ever built assumes the slots are already there, waiting to be contested.

That axiom is breaking. Pew Research Center analyzed the real browsing behavior of 900 U.S. adults during March 2025, covering 68,879 unique Google searches. When a search page displayed an AI summary, users clicked a traditional result link in just 8% of visits, versus 15% on pages without one. Clicks on links inside the AI summary itself? One percent. And 26% of users ended their browsing session entirely after seeing a summary — compared with 16% on pages without one.

In other words, users no longer scan a results page. They finish exploring and comparing inside a single generated answer. The platforms have noticed. OpenAI launched self-serve advertising in ChatGPT in July 2026, with Best Buy, Lowe's, and VistaPrint among the first advertisers, and Google has folded ads into AI Overviews. Both chose the same path of least resistance: transplant the sponsored card from the search era into the AI answer.

But that is a stopgap. The real structural problem is this: when an answer is generated word by word, the advertising opportunity itself becomes a product of the generation process. A research team from Renmin University of China's Gaoling School of AI and Stanford University pushed the question to its logical endpoint. Their mechanism, called LAMA, embeds the ad auction into the token-level generation of the model itself. This is not just a paper trick — it is a roadmap. The object of the next auction is not layout. It is the direction of the answer.

Why "Inserting Ads" Was Never Enough

Today's AI advertising is, at its core, "generate first, paste a card later." The answer is the answer; the ad is the ad; neither touches the other. This inherits the old logic: a finished content container exists, and commercial information gets placed into it.

But a generated answer is fundamentally different from a results page. The page layout is fixed — the third result is always the third result. In an answer, whether a brand appears, where it appears, and how it is introduced all depend on what has already been written. When recommending vacation destinations, discussing transportation before lodging gives different brands completely different positions in the text that follows. Generation doesn't just fill ad opportunities — generation shapes them.

This means advertisers are no longer competing only for the final exposure slot. They are competing for influence over "which direction the answer goes next." The moment influence becomes an object of competition, it needs rules: Who allocates that influence? What happens when an advertiser changes strategy mid-generation? And does the answer stay honest to the user's question, or bend toward the highest bidder?

Existing research — paragraph-level insertion, bidding over candidate answers — splits generation and allocation into two sequential stages. LAMA, short for Latent Advertiser Mixture Auction, fuses them into one process. As the model generates each token, advertisers simultaneously report "how valuable is it to me if the answer continues this way." The platform updates each advertiser's probability of winning final exposure as the text unfolds, and settles once the answer completes.

Generation as Allocation: A Reusable Framework

Abstracted away from the paper, this is a framework I'd call Generation as Allocation: when content production and the allocation of commercial value happen inside the same dynamic process, every static, after-the-fact allocation mechanism fails.

The framework stands on three load-bearing points, each corresponding to a hard requirement in mechanism design:

First, honesty must be an anytime strategy, not an opening posture. LAMA satisfies Markov dominant-strategy incentive compatibility (Markov DSIC): at any reachable state of the generation, an advertiser who keeps reporting truthfully never does worse than one who switches strategies midway. A local-report verification rule requires each advertiser's reports to stay internally consistent, while payments adjust dynamically with the generation — pay when the trajectory turns favorable, receive subsidies when it turns against you. The experiments confirmed the theory: inflating your report does buy more token share and higher allocation probability, but true expected utility peaks exactly at honest reporting. Lying wins influence and loses money.

Second, commercial value cannot override answer quality. LAMA's objective is ad value minus a penalty for deviating from the native answer, with a provable bound on the gap to the ideal mechanism. This is not a moral gesture; it is a survival condition for the product. An ad system that collapses user trust does not get a second chance.

Third, the mechanism must explain real competition. In the sample query "best beach vacations in the world," Expedia appeared early in the answer — essentially a brand anchor used to introduce destinations like the Maldives, Bali, and the Caribbean. Tripadvisor appeared later, with sentences that explained what the product actually does: compare destinations, read traveler reviews, filter lodging and experiences. The platform's probability updates captured this difference and awarded final exposure to Tripadvisor, whose value expression was more complete. What got allocated was not the highest bid, but the value that genuinely held within the text.

The proof of concept used 1,239 real commercial search queries from the Webis Generated Native Ads 2024 dataset, spanning fitness, vacation, and automotive scenarios, each a three-advertiser contest, with Qwen3-14B as the reference model. The result: platform revenue up 10.7% and advertiser value up 3.8%, with no sacrifice on user-experience metrics including relevance, informativeness, and ad naturalness.

Who Gets Repriced by This Framework

Generation as Allocation applies well beyond advertising. Anywhere influence enters a generation process and the generated output determines who wins value, the same logic will reprice everything.

Agent-mediated procurement is the most immediate next battleground. When enterprises let AI agents choose paid tools for their daily workflows, the agent's decision process is itself a generation — and which vendor's information enters that decision context, at what weight, becomes the new "paid ranking." The SEO industry is already migrating toward GEO (generative engine optimization); Digiday reports agencies scrambling to build capabilities that make brands legible to AI systems. All of it is early positioning for influence that enters the generation process.

Push one layer further out: any AI system that generates content in real time and allocates value along the way — personalized feeds, AI customer-service resolutions, automated hiring screens — faces the same mechanism-design problem. They share one trait: there is no "layout" to auction, only a "process" to govern. Auctions used to decide who gets an existing slot. Now the mechanism decides how the slot forms during generation.

There is an uncomfortable symmetry to note. OpenAI's current product design deliberately separates ads from answers — ads are clearly labeled, never alter the model's output, and advertisers receive no access to chats. That is prudence, and it is temporary. Once a mechanism has demonstrated that in-generation allocation lifts revenue without hurting experience, commercial pressure will push platforms from separation toward fusion. Neither regulators nor user expectations are ready for that fusion.

What to Do Now, Depending on Who You Are

  • If you are an advertiser or brand: Stop optimizing only for landing-page conversion. Start optimizing your value expression inside answers. Tripadvisor beat Expedia not on bid size but because its sentences directly answered "what does this product do for me." Being quotable in ad-free organic answers — with real informational density — is the entry ticket to future auctions.
  • If you are a marketing agency: Upgrade GEO from "structured content that AI can parse" to mechanism-level participation — understand what auto-bidding becomes in generative settings and how to design value reports. This is the natural extension of search campaign management, and the window will not stay open long.
  • If you are an AI platform: The boundary between ads and answers is your trust moat; build it before optimizing revenue. LAMA-style mechanisms work on paper, but productization must answer how users perceive and control commercial influence — put user control and clear labeling ahead of the mechanism, not behind it.
  • If you are a publisher or content site: A 1% click-through rate inside AI summaries means "being cited" is replacing "being clicked" as the unit of exposure. Rebuild your traffic strategy and your negotiation leverage around becoming a source for AI answers, not around ranking positions.

The history of advertising is the history of media form deciding what gets auctioned: newspapers auctioned space, search auctioned rank, feeds auctioned impressions. When the medium becomes a single generated answer, the object of the auction migrates with it. The question is no longer whether AI advertising arrives — it is which layer of generation the auction hides in. LAMA offers the first mechanized answer. It may not be the final form, but the direction is now legible.

Scroll to top