Meta-harness optimization loop wired onto Islo sandboxes. POC: 0/5→5/5 in four proposer steps. Built on islo.dev.
-
Updated
May 5, 2026 - HTML
Meta-harness optimization loop wired onto Islo sandboxes. POC: 0/5→5/5 in four proposer steps. Built on islo.dev.
Lynn: open-source desktop/CLI AI Agent with GUI Session Map, headless workers, realtime voice, long-term memory, Brain V2 routing, local 27B/35B GGUF support, and contract-based Agent Regression Kit.
Observe an agent run and GroundEval drafts the policy and diagram for you, no hand-written policy required, then scores what it checked, what it skipped, and what it wasn't allowed to touch.
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.
AI-operated company. Building agent-friend: universal tool adapter for AI agents. @tool → OpenAI, Claude, Gemini, MCP. Live 24/7 on Twitch.
Project page for Meta-harness on Islo (POC). https://zozo123.github.io/meta-harness-on-islo-page/
A reasoning benchmark runner for comparing LLMs as OpenClaw agents use them. 52 prompts, 3 eval sets, 11 traps, LLM-as-judge, tier-based leaderboard.
Scenario Testing for AI Agents
Vendor-neutral research umbrella for measuring AI plugin, agent, and MCP server quality across CLI runtimes (Claude Code, Gemini CLI, Copilot CLI, Codex CLI).
Transcript-first evaluation tool for comparing coding-agent sessions across Codex, Claude Code, and Pi.
Deterministic execution recording and replay for AI agents — turn real agent runs into replayable regression test fixtures.
开源通用 AI Agent 真实任务评测 · 同 Prompt、客观开奖、评分细则全公开 | Open-source evaluation of general-purpose AI Agents on real-world tasks with verifiable outcomes — by PingWest / 硅星人
ckl's personal benchmark for doc writing, infra code, and paper reading — one-click evaluation of the latest models via TUI coding CLIs with an interactive dashboard.
Trajectory-level evaluation for LLM agents — catch cross-turn defects that naive per-turn LLM-as-judge misses, with honest recall measurement
96K param RWKV-7 that detects non-termination (the Gödel sentence analog) zero-shot across SKI combinatory logic, lambda calculus, and Turing machines
Auto-generate evaluation rubrics from agent audit-log trajectories (PhoneWorld pattern applied to action logs)
🚀 基于Java的开源AI自动化评测框架 / An open source AI automation evaluation framework based on Java
An AI coding benchmark task evaluating automated rebase conflict resolution, MLflow tracking security, secret redaction, and Hugging Face model evaluation safety.
PandaProbe Harness turns your agent into a self-healing harness
Add a description, image, and links to the agent-eval topic page so that developers can more easily learn about it.
To associate your repository with the agent-eval topic, visit your repo's landing page and select "manage topics."