Tag: web automation

  • The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

    The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

    Header image source: The Chicago Bears are trending; now comes the hard part • The TRiiBE via The TRiiBE via Google — cropped to 16:9 and colour-adjusted.

    Key takeaways

    • Odysseys tests web agents on complex multi-site tasks humans do daily
    • Best agents score just 1.15% trajectory efficiency—100x slower than humans
    • Agents fail at synthesizing, organizing, and displaying knowledge across sites

    Odysseys was recently released. 200 tasks. Zero mercy. Every single one starts with a Google search page and requires navigating across multiple websites to complete complex objectives. The results? Even the best agents score 1.15% on trajectory efficiency. That’s not a rounding error. That indicates they take many more steps than a human would.

    This isn’t another incremental benchmark. It’s the first one that tests whether web agents can actually work—not just search, but synthesize, organize, and display knowledge across multiple sites. The answer is no.

    The Benchmark Plateau: Why Single-Site Tasks No Longer Cut It

    Those benchmarks are saturated. Frontier models cruise through them with near-perfect scores. But here’s the catch: no one actually needs an agent for that. The real world doesn’t hand you isolated, single-site tasks. It hands you workflows.

    Real web navigation tasks require extended interaction with multiple sites, such as comparing products between retailers or planning travel between booking platforms. Real web navigation tasks require extended interaction with multiple sites, such as comparing products between retailers. Existing benchmarks don’t test any of this. They test whether an agent can click a button. Odysseys tests whether it can think.

    The saturation of single-site benchmarks has created a dangerous illusion. Agents look impressive in demos because the tasks are trivial. But the moment you ask them to do something that requires memory, planning, or cross-domain reasoning, they collapse. That’s not a minor limitation. It’s a fundamental gap between what these systems can do and what they need to do to be useful.

    Introducing Odysseys: A Benchmark for the Hard Part of the Web

    Odysseys isn’t just another dataset. It’s a stress test for the post-search era. Each of the 200 tasks starts from a Google search page and demands navigation across multiple websites to complete a complex objective. Examples:

    • “”
    • “”
    • “”

    These aren’t contrived challenges. They’re the kinds of tasks humans do daily—tasks that require not just finding information, but doing something useful with it. Odysseys isolates this synthesis step. And it’s brutal.

    The tasks are framed as first-person requests, the way a real user might phrase them. That’s intentional. It forces agents to interpret vague, high-level goals—“”—rather than execute rigid, pre-defined steps. And that’s where they fail.

    Most agents handle the first few steps fine. Searching Google. Clicking a link. Extracting some data. But synthesizing that data into a coherent output? Organizing it logically? Formatting it for human consumption? That’s where they fall apart.

    The Efficiency Collapse: Why Agents Take 100x More Steps Than Humans

    Odysseys introduces a metric called Trajectory Efficiency: rubric score per step. Even the best agents score 1.15%. Let that sink in. A human completes these tasks efficiently. An agent takes many more steps.

    This isn’t just a UX problem. It’s a fundamental limitation in planning and memory. Agents can’t:

    • Remember prior actions. They revisit the same page multiple times because they forgot they’d already been there.
    • Plan hierarchically. They don’t break tasks into subtasks or execute them in a logical order.
    • Optimize trajectories. They don’t group related actions to minimize steps.

    The consequences are severe:

    • Cost. More steps mean more API calls. Higher operational costs.
    • Latency. Users won’t wait for an agent that takes 10 minutes to complete a task they could do in 30 seconds.
    • Brittleness. More steps mean more failure points. A single misclick can derail the entire workflow.

    This is why the hype around “autonomous agents” feels premature. If they can’t match human efficiency, they’re not autonomous. They’re just expensive macros.

    The Rankings: Who Won, Who Lost, and Why It’s Not Close

    Nine agents were evaluated on Odysseys: frontier models, open-weight models, and terminal-style agents. The metrics:

    • Rubric average. How well they completed tasks.
    • Perfect-task rate. How often they completed tasks flawlessly.
    • Trajectory efficiency. That 1.15% metric.
    • Trajectory-level LLM judge. A secondary evaluation of action quality.

    The results reveal a steep drop-off after the top performers. Most agents failed to complete many tasks perfectly. Here’s the breakdown:

    Frontier Models

    The best performers. Still struggling. Their strength? Raw reasoning power. Their weakness? Long-horizon coherence. They’d often forget intermediate steps or fail to synthesize data from multiple sources. Agents often struggle to synthesize data from multiple sources into coherent outputs.

    Open-Weight Models

    Showed promise in niche tasks. Lacked consistency. Their issue? Limited context windows and weaker reasoning. They’d excel at tasks with clear, step-by-step instructions but collapse on open-ended workflows.

    Terminal-Style Agents

    Often failed due to rigid action spaces. These agents are designed for command-line interactions, not dynamic web UIs. They’d get stuck on pages with complex layouts or fail to handle UI changes—a pop-up modal blocking an action, for instance.

    The rankings read like a report card where the valedictorian gets a C-. That’s not progress. It’s a wake-up call.

    The Synthesis Gap: Why Agents Can’t Turn Information Into Knowledge

    Odysseys tasks require agents to synthesize, organize, and display information. This is where they fail most spectacularly. Let’s break it down:

    Synthesis

    Agents struggle to merge data from multiple sources. Example: Comparing laptops from three retailers. The specs might be formatted differently—“16GB RAM” vs. “Memory: 16GB.” Agents often fail to normalize this data, leading to incomplete or inconsistent outputs.

    Organization

    Agents lack hierarchical reasoning. They’ll dump raw data into a document without grouping related items or prioritizing information. Agents lack hierarchical reasoning for organizing information.

    Display

    Agents treat output as an afterthought. They’ll paste unstructured text into a document instead of formatting it for readability. Agents treat output as an afterthought rather than formatting it for readability.

    A human would:

    1. Extract key specs—range, price, charging time—from each site.
    2. Organize them into a consistent format.
    3. Design slides with clear headings, bullet points, and visuals.
    4. Save the file in the correct format.

    An agent might:

    1. Extract specs but miss key details—ignoring charging time, for instance.
    2. Paste them into a slide as unformatted text.
    3. Fail to save the file correctly—saving as a.txt instead of.pptx.

    This isn’t a minor flaw. It’s a fundamental limitation in how agents process and output information. Humans don’t just collect data. We curate it. Agents don’t.

    The Accessibility Angle: Why This Matters Beyond Tech Demos

    Advancing autonomous web agents isn’t just about productivity. It’s about accessibility. For users with disabilities or limited technical skills, these tools could be transformative.

    Current state: Assistive tools—screen readers, for example—help with navigation but not synthesis. They can read a webpage aloud, but they can’t compile a research report or plan a trip across multiple sites.

    Future potential: Agents that handle long-horizon tasks could empower users to complete complex workflows without manual intervention. Example: A user with motor impairments could ask an agent to handle complex multi-site workflows.

    Barrier: The efficiency and reliability gaps in Odysseys suggest these tools aren’t yet viable. If an agent takes many more steps than a human, it’s not accessible. It’s a frustration.

    This is the most compelling argument for solving these problems. If agents can’t handle Odysseys tasks, they’re not helping the people who need them most.

    The Path Forward: Reinforcement Learning, Inference-Time Search, and Beyond

    The Odysseys paper suggests two promising directions for improving agents: reinforcement learning (RL) and inference-time search.

    Reinforcement Learning

    RL could help agents learn efficient trajectories through trial and error. The idea: reward agents for completing tasks with fewer steps. Over time, they’d optimize their workflows.

    Problem: Sample inefficiency. RL requires massive amounts of data. Web interactions are expensive—both in time and cost. Training an agent to plan a trip might require millions of simulated interactions, each with its own API calls and latency.

    Inference-Time Search

    Techniques like Monte Carlo Tree Search (MCTS) could improve planning depth. Instead of executing actions immediately, agents could simulate possible trajectories and choose the most efficient path.

    Problem: Latency. MCTS is computationally intensive. Real-time agents can’t afford to spend seconds deliberating over every action.

    Alternative Approaches

    Hybrid models—combining LLMs with symbolic planners—might bridge the gap. Example: An LLM could generate a high-level plan while a symbolic planner handles low-level execution. ”—while a symbolic planner handles the low-level execution.

    Opinion: The most promising direction might be teaching agents to “” before acting. Humans don’t execute tasks blindly. We break them into subtasks, anticipate obstacles, adapt. Agents need to do the same.

    The Big Picture: What Odysseys Tells Us About the State of AI

    Odysseys exposes a critical truth: search was the easy part. The real challenge is what comes after—turning raw information into actionable knowledge.

    Prior benchmarks measured whether agents could find information. Odysseys measures whether they can use it. The gap between the two is vast.

    Here’s the contrast:

    • Single-site tasks: “”
    • Odysseys: “”

    The answer, for now, is no. But Odysseys gives us a way to measure—and eventually close—that gap.

    The most interesting open question isn’t whether agents can pass Odysseys. It’s whether the techniques that help them pass—RL, inference-time search, hybrid planning—will generalize to other domains. If they do, we might finally see agents that are truly useful. If they don’t, we’re stuck with expensive macros that can’t handle the hard work of knowledge synthesis.

    Either way, Odysseys has redefined how we evaluate agents. And that’s progress.