LoHoSearch – The Next-Generation Search Intelligence Agent Evaluation Benchmark Launched by Meituan

LoHoSearch – The Next-Generation Search Intelligence Agent Evaluation Benchmark Launched by Meituan

What is LoHoSearch

LoHoSearch is the next-generation search agent evaluation benchmark launched by Meituan LongCat, automatically generating questions based on a knowledge graph encompassing 7.62 million Wikipedia entities and 265 million relation edges, replacing traditional manual question creation. The benchmark includes 544 high-quality questions manually verified across 11 domains including music, film, and geography, designed to challenge models through controlled dimensions of search space size and structural complexity. The current strongest model, GPT-5.5, achieves only 34.74% accuracy—far below its over 90% performance on BrowseComp—providing a more discriminative evaluation standard for long-horizon search reasoning and context management.

LoHoSearch

Main Features of LoHoSearch

  • Automated Knowledge Graph Construction: Utilizing the full English Wikipedia as data source, it constructs a hyper-scale knowledge graph with 7.62 million entity nodes and 265 million directed edges, each entity annotated with Wikidata type and in-degree popularity, laying the foundation for global perspective-based question selection.
  • Two-Dimensional Difficulty Control: Systematically modulates question difficulty from both "search space" and "structural complexity" orthogonal dimensions, surpassing cognitive limits inherent in manual question design.
  • Tree and Graph Structure Sampling: Tree structures amplify search space to increase difficulty; graph structures introduce cyclic dependencies and cross-constraints atop vast search spaces, further compounding structural complexity.
  • Automated QA Generation and Validation: Extracts Wikipedia descriptions from subgraphs, rewrites them into natural language questions via large models, then applies three-stage automated filtering—relation extraction, subgraph coverage check, answer satisfaction verification—followed by human review to ensure quality.
  • Multi-Model Performance Benchmarking: Supports standardized evaluation of mainstream search agents, outputting core metrics such as accuracy, tool invocation count, and calibration error to clearly reflect long-horizon reasoning capabilities.

Technical Principles of LoHoSearch

  • Automated Knowledge Graph Construction: Uses the complete English Wikipedia as input, extracting 7.62 million entity pages as nodes and 265 million hyperlinks from article bodies as directed edges. Each entity is annotated with Wikidata P31 type and in-degree popularity, forming a global knowledge network.
  • Two-Dimensional Difficulty Definition: Quantifies difficulty along two orthogonal axes—"search space" and "structural complexity"—breaking the cognitive constraints of manual question design.
  • Tree Structure Subgraph Sampling: Increases difficulty by expanding the search space; from a single answer node, multiple independent branches are expanded downward, each corresponding to a high-candidate constraint, causing elimination cost to grow exponentially.
  • Graph Structure Subgraph Sampling: Introduces cyclic dependencies and cross-constraints on top of massive search spaces; multiple condition nodes interconnect, requiring simultaneous satisfaction of multiple interlocking relationships, creating dual-layer challenges through compounded structural complexity.
  • Relation Extraction and Natural Language Generation: Extracts Wikipedia descriptions from sampled subgraphs and uses large models to rewrite them into coherent natural language questions, ensuring all clues preserve original relational information within the subgraph.

How to Use LoHoSearch

  • Load Data: Directly load meituan-longcat/LoHoSearch via the datasets library to access 544 questions and ground-truth answers.
  • Inspect Structure: The dataset contains question text, corresponding subgraph structure, domain labels, and uniquely validated answers verified against the knowledge graph.
  • Integrate for Evaluation: Feed questions into the target search agent and record its reasoning trajectory and final output.
  • Automated Scoring: Compare model outputs against the provided ground truth to verify compliance with all constraints and compute accuracy.
  • Comparative Analysis: Use attached subgraph complexity metrics and domain labels to analyze model performance across different difficulty levels and thematic areas.

Core Advantages of LoHoSearch

  • Breakthrough Beyond Manual Questioning Limits: Traditional benchmarks are constrained by annotators’ known entity scope. LoHoSearch leverages a global knowledge graph to systematically identify “true hard problems” with high search space and high structural complexity.
  • Significantly Higher Difficulty Than Existing Benchmarks: Same model achieves 58.84% accuracy on BrowseComp but only 10.02% on LoHoSearch, with median tool calls increasing from 35 to 61—a 74% rise in challenge intensity.
  • Superior Discriminative Power: Context management strategies improve performance by 14.03% on BrowseComp but only 6.8% on LoHoSearch, better reflecting their marginal gains in long-horizon complex scenarios.
  • Independently Controllable Structural Complexity: Graph-structured question accuracy (8.01%) is significantly lower than tree-structured ones (11.89%), proving structural complexity is an independent difficulty factor beyond search space, quantifiable in isolation.

Project Repository for LoHoSearch

  • HuggingFace Dataset Hub: https://huggingface.co/datasets/meituan-longcat/LoHoSearch
  • arXiv Technical Paper: https://arxiv.org/pdf/2606.12837

Competitive Comparison with Similar Benchmarks

Dimension LoHoSearch BrowseComp
Question Generation Method Knowledge Graph Auto-Generation + Human Verification Manual Design
Number of Questions 544 Hundreds
Domain Coverage 11 domains (music, film, geography, etc.) Not explicitly domain-segmented
Difficulty Control Two-dimensional: Search Space + Structural Complexity Subjective, human-controlled
Top Model Accuracy GPT-5.5: 34.74% GPT-5.5 Pro: 90.1%
Median Tool Invocations 59 26
Context Strategy Improvement +6.8% +14.03%
Saturation Level Far from saturation, high discriminative power Approaching saturation
Use Case Long-horizon search reasoning, context management evaluation Basic search capability evaluation

Application Scenarios of LoHoSearch

  • Search Agent Capability Benchmarking: Provides a high-difficulty, standardized test for search agents with long-horizon reasoning, enabling precise identification of weaknesses in complex information retrieval chains.
  • Context Management Strategy Validation: Evaluates real-world gains of context compression and validation strategies—such as Summary, Discard-all, Verify—in ultra-long interaction scenarios, preventing overestimation of strategy effectiveness on simple benchmarks.
  • Model Iteration Benchmarking: Offers an unsaturated evaluation standard for research teams, supporting capability boundary probing and performance comparison during model version iteration.
  • Search Algorithm Research: Supplies a controllable-difficulty dataset and evaluation framework for complex search algorithm research requiring multi-step reasoning and cross-validation.
  • Agent Product Selection: Assists enterprises in selecting search agent solutions by distinguishing actual model performance under real-world complex query scenarios using LoHoSearch.
#AI Tools

Disclaimer: Contains third-party opinions, does not constitute financial advice

Share To
X
Telegram
WeChat
QQ
Link
Recommended Reading

Wan-Streamer v0.2 – A Multimodal Understanding and Generation Model Launched by Alibaba Tongyi

4 days ago
Wan-Streamer v0.2 – A Multimodal Understanding and Generation Model Launched by Alibaba Tongyi

StaffDeck – Open-source Enterprise-grade Digital Employee Platforms by Membrane Intelligence and Others

4 days ago
StaffDeck – Open-source Enterprise-grade Digital Employee Platforms by Membrane Intelligence and Others

Nemotron 3 Embed – NVIDIA's Open-Source Text Embedding Model Series

4 days ago
Nemotron 3 Embed – NVIDIA's Open-Source Text Embedding Model Series

Inkling – A multimodal foundational model launched by Thinking Machines Lab

5 days ago
Inkling – A multimodal foundational model launched by Thinking Machines Lab

Backup – K12 Primary and Secondary School Teachers' AI Lesson Planning Platform Launched by 360

5 days ago
Backup – K12 Primary and Secondary School Teachers' AI Lesson Planning Platform Launched by 360

X2.0 – The World's First Real-Time Interactive Video Generation Model Launches by Xmax AI

6 days ago
X2.0 – The World's First Real-Time Interactive Video Generation Model Launches by Xmax AI

Xiaomi-Robotics-U0 – Xiaomi's Unified Embodied Synthetic Model

6 days ago
Xiaomi-Robotics-U0 – Xiaomi's Unified Embodied Synthetic Model