Challenges4 000CP pool
ClosedCode

Determinist Evaluator Grid

Build the evaluation backbone: structured JSON grids (criteria, sub-criteria, scoring anchors) and the AI agent that scores submissions against them. The goal is reproducibility. The same delivery must get the same score, verified by repeated test runs before a grid goes live.

2 participants · own branch each

Why this challenge exists

Every evaluation on the platform (code challenges, ML submissions, sandboxes) goes through one evaluate capability. It scores a bundle against a structured grid. For scores to be credible, including to clinical partners, the same delivery must get the same score.

Who hosts this challenge

MyTwin Lab

The brief

Objective

Grids live in the database (code, dataset, model), with criteria scored 0–9 and weights summing to ~1. Admins edit, draft and archive them from /admin/evaluation-grids. A test run scores a sample five times in a row to check a grid's consistency before it goes live.

Expected result

• A missing grid fails before any GitHub or Kaggle call, with no fallback to a frozen copy

• The bundle is written to a confined temporary workspace and cleaned up whether the run succeeds or fails

• Every run is recorded in evaluation_runs (grid, source, duration, raw score), and failed runs can be retried