Research · September 21, 2026

Pushing the Capabilities of Computer Use Agents with Intent Detection

Introduction

Browser agents are getting remarkably good at execution, but they have a design flaw: prompting them often takes longer than doing the job yourself. We’re sharing an early look at some of our research that explores predicting your next action without explicit instruction. If you are trying to book a flight, writing a prompt that specifies your preferred dates, layovers, budget and other minor details takes genuine effort. Browser interactions also raise the stakes, since many actions are irreversible, one wrong click can trigger a non-refundable charge or send an email you can’t take back..

Removing the Prompt

Standard browser agents require instructions, page context, and a browsing history. We chose to drop the instruction, inferring intent from the page and history alone.

While the page can show you what actions are possible, your history gives the intent that drives your next action. If you open Google Flights after reading an email with travel dates, the right move is typing them in. A confirmation receipt looks like a dead end, unless a teammate is waiting on those details.

This is also why we explored predicting one action at a time rather than a whole plan.The correct action is inherently dependent upon what each page shows , so instead of committing to a script it can’t revise, the model predicts an action, sees the resulting page, and decides again. Each step gives it fresh evidence and keeps you in the loop. Making that work in real time came down to a few technical hurdles.

Our approach

Our research is post-trained on Qwen3.8-27B. We fine-tune it on economically valuable browser sessions, then run a round of reinforcement learning. The entire prediction happens in a single forward pass, because the prediction has to arrive before the user acts.

What a training example looks like. Each training row has the same shape as a production request, containing the previous actions and a custom DOM representation. This is essentially a compact text representation of the page that is smaller than the raw DOM but richer than a bare list of interactive elements. The ground truth label is the next action in the format of one of the following:

  • Clicking on a target element
  • Typing a value into a target element field
  • Abstaining (or “no-op”) when uncertain

We deliberately left out the cursor position, which heavily correlates with action selection in human session recordings. Since our users navigate via keyboard and tab through suggestions, training on cursor data would give the model a false signal that would not exist in the new paradigm.

A dense next-action policy. We fine-tune with a BF16 LoRA adapter, rank 64, and alpha 128, over a 16,384-token context for two epochs on four B200s on roughly 16.5K examples (about 110M tokens), training only on the final action. We used a large adapter so that our SFT teaches the model to infer intent, a net-new capability, rather than simply fitting on the output format. We also saw that data cleanliness mattered a lot. When the user’s past actions clearly explained their next step, intent inference emerged on its own, without having to explicitly train for it.

Selective abstention. An instructed agent is expected to act; a proactive one must know when to act. An unprompted, incorrect browser click causes far more friction than doing nothing. On the other hand, excessive abstention disrupts the workflow’s momentum. So, we track two quantities: when the model acts correctly (accuracy), and the fraction of opportunities on which it acts (coverage).

Supervised training alone taught the model that no-op was permitted (making up 3% of rows in SFT), but the resulting policy still acted almost everywhere, reaching 49.71% element accuracy at 97.20% coverage. We then trained abstention with reinforcement learning. Every prompt had a known correct action, where no-op was a safe fallback rather than a positive objective:

+0.75  correct method and element
 0.00  no-op
-0.25  wrong action

We generated sixteen rollouts per prompt and compared two RL objectives. The first scores each rollout on its own reward against a fixed baseline (an absolute-reward, REINFORCE-style policy gradient), so a prompt whose sixteen rollouts are all wrong still pushes the model down. The second scores each rollout against the group’s mean reward (a group-relative advantage with no standard-deviation normalization, the Dr. GRPO variant, using a GSPO sequence-level clipped importance ratio and no KL). It is sharper but has two quirks: if every rollout in a group earns the same reward there is nothing to compare and the group teaches nothing, and when most rollouts are wrong a no-op scores 0 and looks good next to them. So the model can learn to stay quiet just because the alternatives were strictly worse. Our best run combined the two ideas: the absolute-reward objective first, then the group-relative one.

Right more often, acting less often RL step: early to late SFT baseline Hover or tab to a checkpoint for its accuracy and coverage 50% 55% 60% 65% 70% 75% 40% 50% 60% 70% 80% 90% 100% Coverage: how often the model acts Element accuracy when acting SFT baseline SFT baseline · 49.71% accuracy when acting · 97.20% coverage ckpt 128 · 56.48% accuracy when acting · 87.93% coverage ckpt 256 · 59.36% accuracy when acting · 79.07% coverage ckpt 384 · 63.64% accuracy when acting · 66.73% coverage ckpt 512 · 63.67% accuracy when acting · 65.33% coverage ckpt 640 · 67.52% accuracy when acting · 54.80% coverage ckpt 768 · 61.43% accuracy when acting · 73.47% coverage ckpt 896 · 69.20% accuracy when acting · 39.40% coverage ckpt 1024 ckpt 1024 · 72.95% accuracy when acting · 35.73% coverage ckpt 1152 · 65.61% accuracy when acting · 58.93% coverage ckpt 1280 · 65.04% accuracy when acting · 61.40% coverage ckpt 1408 · 67.46% accuracy when acting · 53.27% coverage ckpt 1536 · 66.94% accuracy when acting · 56.87% coverage ckpt 1664 · 66.34% accuracy when acting · 54.07% coverage ckpt 1682 · 67.19% accuracy when acting · 51.20% coverage

RL let us trade coverage for accuracy. The more we trained, the more cautious the model got. It acted less often, but was right more often when it did. At its most cautious checkpoint (ckpt 1024), the model acted about 36% of the time and was right 73% of the time it acted. Coverage fell, bottomed out, then climbed back up in later checkpoints as a result of the group-relative advantage paradigm. When the model no-ops on almost everything, most rollouts in a group get no reward, so a single correct answer lands well above the group mean and gets a large positive advantage. This nudged the model to act, so coverage recovered instead of collapsing to zero.

Reinforcement learning mostly converted wrong actions into no-ops; it did not teach the model any new correct actions. It worked only because the supervised policy was already good enough to land some correct actions among its rollouts, which gave the optimizer a usable preference. An earlier attempt on a weaker policy had nothing to reward.

Low-latency serving. To beat human motor reaction time, we had to pull inference from a couple of seconds down to under 400 milliseconds. DFlash tuning did most of the heavy lifting, alongside capping the token budgets. Experimentations with quantization lowered accuracy to unacceptable levels.

What we found

The base model is important. We began on Qwen3.5-A3B, a mixture-of-experts model with about 3B active parameters. There we midtrained on descriptions of custom DOM representations to teach the model about sites it had never seen in pretraining: internal dashboards, monitoring tools, admin panels with ambiguous labels. This raised accuracy, but its most useful effect was on model confidence, which is needed for the model to decide when to stay quiet. Moving the base model to Qwen3.8-27B removed the need for midtraining. We suspect that it had already seen a lot of computer-use data in pretraining, so it was familiar with web pages and no longer needed the midtraining step. Being a dense model also helps as it uses far more active parameters for each decision.

The changing element IDs weren’t the bottleneck. We assumed the hard part would be selecting the right DOM element IDs which are volatile, with selectors changing across every page and reload. To test that, we built a pointer model that sidesteps IDs entirely: instead of writing out an ID, it turns every element on the page into a vector, scores them, and picks the highest-scoring one with a softmax over the elements. It could reach almost any target yet still lost to plain element ID generation. Graph message passing, depth embeddings, and hard negatives didn’t help. The errors gave it away; they weren’t near-misses. The model kept choosing elements that moved the user towards a completely different goal, so it wasn’t getting the ID wrong, it was getting the intent wrong.

Representation was not the problem. The tokenizer writes numbers digit by digit, so the model builds an ID like 0-194 from the digits 0 - 1 9 4 instead of memorizing it, and a page of unseen IDs is just familiar digits in a new order. IDs are also numbered in tree order, so close IDs sit together on the page and a near-miss lands on a neighbor.

The bottleneck is inferring intent. Intent is hierarchical. The goal (send the flight times to a colleague) persists across Gmail, Google Flights, and the airline site, while the immediate subgoal changes on every page. History explains why the user is there; the page constrains what is possible next.

Feeding the model a longer history gave surprisingly little lift, but our context experiments established two distinct facts. First, chronological order carries real signal: an ordered 128-state window outperforms a shuffled baseline. Second, recency dominates: a 128-state window barely outperforms a 25-state window. Because almost all of the signal lives in the last few actions, we swept the window from 25 down to 5 and settled on 15. At 25 we spent context on stale actions no longer relevant to the task, and below 15 we dropped actions that still mattered.

We also tried to supervise intent explicitly in a few different ways.

First, we ran a supervised stage to explicitly learn the user’s goal. It predicted the overall intent correctly about 54% of the time, but initializing the action model from that checkpoint did worse than initializing from the base model. The training data was the problem. The trajectories were human, but the intent was LLM-labeled which overweighted incidental details and dropped the ones that drove the choice. A label might fixate on the exact 2:15 PM departure the user selected while missing the budget constraint influencing that decision. With these labels, the model learned to recover the broad goal but latched onto irrelevant details that hurt the action model.

Alternatively, we tried giving the model reasoning tokens before the predicted action. Its logic was usually sound but built on the wrong intent, so it reasoned cleanly to a wrong action. This also increased latency.

The clearest finding came from factoring the task into two stages: first predicting the desired effect on the page, then the action to produce it. The model frequently predicted the wrong effect, but the factorization cleanly localized the issue. When conditioned on the correct effect, it selected the right element almost every time. Almost all the difficulty lies in deciding what should happen next, not in grounding the decision to an element. Once the model knows what it wants, choosing the action is easy.

What’s next

We’re now looking at how far this generalizes beyond the sites and workflows in our training data, and at what changes when predictions start chaining into longer sequences. More detail in a follow-up post later this week.