The brief
Objective
Grids live in the database (code, dataset, model), with criteria scored 0–9 and weights summing to ~1. Admins edit, draft and archive them from /admin/evaluation-grids. A test run scores a sample five times in a row to check a grid's consistency before it goes live.
Expected result
• A missing grid fails before any GitHub or Kaggle call, with no fallback to a frozen copy
• The bundle is written to a confined temporary workspace and cleaned up whether the run succeeds or fails
• Every run is recorded in evaluation_runs (grid, source, duration, raw score), and failed runs can be retried