Skip to content

Test Case Groups

A test-case group is a repo-defined, ordered set of related test cases, such as the tower-defense cases. Groups exist for presentation: the home page renders one cross-case leaderboard per group, so a visitor sees which models build the best game of each kind without opening every member case.

Groups are global and shared by every visitor, and they decide only what the home page presents. A case’s tags classify one case for filtering, and the per-account case groups of coverage plans schedule a reviewer’s runs.

A group is a directory under test-case-groups/<slug>/ in the repository, containing one manifest, test-case-group.toml, which declares:

  • slug, the stable identifier, matching the directory name;
  • name, the display name;
  • summary, an optional one-line description;
  • rank, an optional ordering key. Groups order by rank ascending, then by name, and ranked groups precede unranked ones.
  • cases, the ordered list of member test-case or game-jam slugs. The list must be non-empty and free of duplicates.

A manifest carrying an unrecognized field is rejected. The catalogue is read from the checkout at run time, like the test-case catalog, so authoring a group is a manifest edit with no rebuild. tcab test-case-groups lists the catalogue.

Every member slug must resolve in the test-case catalog, which covers test cases and game jams alike. The repository’s manifest test checks this against the checkout, and the backend’s ingest checks it again against the catalog it is ingesting: a group naming a slug the catalog cannot resolve is rejected with a logged error, and the remaining groups still ingest.

The backend parses the groups on a whole-catalog ingest scan and writes the set to its definition store, so a deployment serves the same groups a local checkout resolves. The set is served in display order at GET /test-case-groups, with the ordering rank already applied. An ingest that changes the set queues a public snapshot refresh, and the snapshot carries the set as its own object, so the static gallery renders the same groups; see Public Snapshot.

A group’s leaderboard is folded from the scored runs of its member cases at each case’s current version, grouped by harness, model, and engine. Entries rank by mean score fraction: a run contributes its earned share of its own case’s checklist weight, so runs of cases with different point totals stay comparable. Ties break by the better best rating, then by recency. A game jam member contributes its whole-game grade in place of a functional rating.