Skip to content

Actions: benchflow-ai/awesome-evals

Actions

All workflows

Actions

Loading...
Loading

Showing runs from all workflows
246 workflow runs
246 workflow runs

Filter by Workflow

Filter by Event

Filter by Status

Filter by Branch

Filter by Actor

Add PerspectiveGap and Universal Magic Words
Awesome-Evals Responder #123: Issue comment #43 (comment) created by xdotli
2s
Add agent-qa
Awesome-Evals Responder #119: Issue comment #36 (comment) created by xdotli
9s
Add Prompt-TDD Methodology to §5a
Awesome-Evals Responder #118: Issue comment #40 (comment) created by xdotli
1s
Add verdict4 to 5a (eval frameworks & harnesses)
Awesome-Evals Responder #117: Issue comment #49 (comment) created by xdotli
2s
Add FinMirror paired-world RAG evaluation
Awesome-Evals Responder #116: Issue comment #58 (comment) created by xdotli
10s
Add Simplified-Chinese LLM response QA toolkit
Awesome-Evals Responder #114: Issue comment #65 (comment) created by xdotli
1s
Add ClawBench to agent-specific evaluation
Awesome-Evals Responder #112: Issue comment #54 (comment) created by xdotli
1s
Add Confident AI to 5f (observability + eval platforms)
Awesome-Evals Responder #111: Issue comment #62 (comment) created by xdotli
11s
Add StructEval benchmark
Awesome-Evals Responder #110: Issue comment #57 (comment) created by xdotli
1s
Add greenproof to §8 (verifiers)
Awesome-Evals Responder #109: Issue comment #50 (comment) created by xdotli
1s
Add Coder Eval to 5a · Eval frameworks & harnesses
Awesome-Evals Responder #106: Issue comment #61 (comment) created by xdotli
Skipped
Add ClawBench to agent benchmark resources
Awesome-Evals Responder #105: Issue comment #64 (comment) created by xdotli
1s
Make the daily scan's "no new items" verifiable
Awesome-Evals Responder #103: Pull request #66 opened by xdotli
2m 26s
Eval Scan (audit + update)
Eval Scan (audit + update) #52: Scheduled
2m 41s main
Add Simplified-Chinese LLM response QA toolkit
Awesome-Evals Responder #102: Pull request #65 opened by 10Lollipop
13s
Add ClawBench to agent benchmark resources
Awesome-Evals Responder #101: Pull request #64 opened by reacher-z
23s
Eval Scan (audit + update)
Eval Scan (audit + update) #51: Scheduled
2m 29s main