Vibe Bench
Benchmarks evaluating long-horizon tasks that require sustained human-AI interaction
Pinned Loading
Repositories
Showing 5 of 5 repositories
- VibeLifeBench_livedemo Public
- VibeLifeBench Public Forked from evolvent-ai/VibeLifeBench
🗓️ The hardest life-admin benchmark for agents — lawsuits, escrow shortfalls, apartment hunts, exams. 20 long-horizon tasks × 20–30 stages across 10 domains and 21 services, scored by 1247 atomic checks that read backend state, not prose. Bilingual zh/en.
- VibeLifeBench_homepage Public
- VibeSearchBench.github.io Public
- VibeSearchBench Public
🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored by verifiable schema-free knowledge-graph evaluation. No vibes, just triplet F1.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…