invariant systems / blog

Announcing Windtrap

2026-09-27 · Thibaut Mattio

At Invariant, we're exploring how we can make code generated by agents safer and more trustworthy.

We're exploring this question from several angles, including building a specialized coding harness (mentat), improving the agent feedback loop through static verification and linting (litany), and more.

One high-leverage axis we're working on is the idea that agents don't work in a vacuum: they tend to replicate what's in your repository, from the tone of voice in your documentation to the structure and coding patterns you use. They add to what is already in place rather than do whatever they would do from a blank slate.

This creates a snowball effect: if your code is messy and full of hacks and anti-patterns, the agents will multiply that. On the other hand, if you have great documentation and tidy implementations and interfaces, the output of your agents will be of much higher quality.

This is a very useful property to leverage when trying to get an agent to write trustworthy code. For instance, if your project is full of benchmarks, you'll see your agent add benchmarks for the work it does without you asking for it. In the long run, the result is a project with a good performance profile, because the agents consistently gated their work on performance regressions.

And so one useful framing of our larger attempt at improving the quality of agent output is: "what is a blueprint for an OCaml project that will make the project higher quality over time if what is in there is multiplied?"

Strong testing hygiene is obviously one of those things, and so we built Windtrap: a testing framework for OCaml, with the goal of giving both humans and agents the best tool possible to test OCaml projects with as little friction as possible. Windtrap test suites should be as exhaustive as possible and read like prose.

We're thankful to be building on the shoulders of giants: the OCaml ecosystem has a very strong culture of testing and correctness, with many testing frameworks and tools, each implementing elegant ideas for finding bugs in your code efficiently. Unfortunately, while deep, the ecosystem is quite sparse. Each testing method or idea is implemented in its own project, often incompatible with the others. This makes building a solid test suite, using the right kind of test for each situation, quite daunting, and you end up with a project that looks like a Frankenstein of testing setups. Not exactly the pristine blueprint that we want our agents to be multiplying.

Windtrap

Windtrap's main contribution is to bring all these testing methods and ideas into one library, with no dependencies beyond OCaml (the optional PPX uses ppxlib), and a beautiful API where every kind of test composes with the others into one cohesive output. It supports:

Here's a small test suite for two modules: Lru, a least-recently-used cache built on a hash table, and Rle, a run-length encoder:

open Windtrap
open Cache

(* A reference for Lru: a list of bindings, the most recently used first. *)
module Model = struct
  type t = { capacity : int; mutable entries : (int * int) list }

  let create capacity = { capacity; entries = [] }
  let length m = List.length m.entries
  let to_list m = m.entries

  let add m k v =
    let rest = List.remove_assoc k m.entries in
    let rest = List.filteri (fun i _ -> i < m.capacity - 1) rest in
    m.entries <- (k, v) :: rest

  let find m k =
    let found = List.assoc_opt k m.entries in
    Option.iter (add m k) found;
    found
end

let cache =
  abstract "c" ~pp:(fun ppf m ->
      Testable.pp (list (pair int int)) ppf m.Model.entries)

let key = Gen.int_range 0 3

let lru =
  group "lru"
    [
      test "evicts the least recently used" (fun () ->
          let c = Lru.create 2 in
          Lru.add c "a" 1;
          Lru.add c "b" 2;
          ignore (Lru.find c "a");
          Lru.add c "c" 3;
          List.iter (fun (k, v) -> Printf.printf "%s=%d\n" k v) (Lru.to_list c);
          expect (output ())
          @@ __POS_OF__ {|
            c=3
            a=1
          |});
      stateful "behaves like a list of bindings"
        [
          command "create"
            (Gen.int_range 1 4 @-> makes cache)
            Model.create Lru.create;
          command "add"
            (cache ^-> key @-> Gen.int_range 0 99 @-> returns unit)
            Model.add Lru.add;
          command "find"
            (cache ^-> key @-> returns (option int))
            Model.find Lru.find;
          command "length" (cache ^-> returns int) Model.length Lru.length;
          command "to_list"
            (cache ^-> returns (list (pair int int)))
            Model.to_list Lru.to_list;
        ];
    ]

let rle =
  group "rle"
    [
      test "encodes runs" (fun () ->
          equal
            (list (pair int char))
            [ (3, 'a'); (1, 'b') ]
            (Rle.encode "aaab"));
      prop "decode inverts encode"
        (Gen.string_of (Gen.char_range 'a' 'c'))
        (Law.round_trip string (list (pair int char)) Rle.encode Rle.decode);
    ]

let () = exit (run "cache" [ lru; rle ])
$ dune exec test/test_cache.exe -- -v
cache: 4 tests (seed s1:891f7248fbd4ad7c)
  PASS  lru › evicts the least recently used       0.2ms
  PASS  lru › behaves like a list of bindings      2.9ms
  PASS  rle › encodes runs                         0.0ms
  PASS  rle › decode inverts encode                0.2ms
4 passed in 3.7ms.

As you can see, unit, expect, property and stateful tests all blend into one cohesive test suite.

The stateful test is where most of the strength lies. We never wrote a test case for it: we described the cache's API once, gave a list as its reference, and Windtrap draws programs of calls and compares every result. Here's what it prints when find forgets to mark a key as recently used:

$ dune exec test/test_cache.exe -- --seed s1:9ea6ce86246615ee -f behaves
cache: 1 test (seed s1:9ea6ce86246615ee)
──────────────────────── failures ────────────────────────
  FAIL  lru › behaves like a list of bindings
    test/test_cache.ml:44
      44 │ stateful "behaves like a list of bindings"

    counterexample (case 15, shrunk 9 steps): 5 calls, last: to_list
       #  reference before  call
       1                    let c1 = create 2
       2  []                add c1 2 0
       3  [(2, 0)]          add c1 0 0
       4  [(0, 0); (2, 0)]  find c1 2
       5  [(2, 0); (0, 0)]  to_list c1
    which failed at:
      test/test_cache.ml:56
        56 │ command "to_list"
      call 5 of 5: to_list c1
      expected  [(2, 0); (0, 0)]
                  ~       ~
      actual    [(0, 0); (2, 0)]
                  ~       ~
──────────────────────────────────────────────────────────

replay: dune exec test/test_cache.exe -- --seed s1:9ea6ce86246615ee -f 'behaves'
1 failed in 2.2ms.

The failing program is shrunk to the five calls that matter, and the replay: line reruns it exactly.

Test coverage takes two commands:

$ dune runtest --instrument-with ppx_windtrap.coverage
cache: 4 passed in 3.8ms (seed s1:d343590d543a1e5f).
$ dune exec windtrap -- coverage -u
   cover    points   file         uncovered lines
   93.8%    30/32    lib/lru.ml   8
  100.0%    12/12    lib/rle.ml

lib/lru.ml: 93.8% (30/32)

      7 │ let create capacity =
  ▌   8 │   if capacity < 1 then invalid_arg "Lru.create: capacity must be positive";
      9 │   { capacity; table = Hashtbl.create capacity; clock = 0 }

coverage: 95.5% (42/44 points)

No test creates a cache with a capacity of zero!

If coverage tells you which code your tests run, mutation testing tells you which of it they actually check: it changes the code one small rewrite at a time and reports each change that no test noticed. It takes one command:

$ WINDTRAP_MUTATE=1 dune runtest --force --instrument-with ppx_windtrap.mutate
cache: 4 passed in 3.9ms (seed s1:9c8104b41db35d5a).

─────────────────────── survivors ────────────────────────
  SURVIVED  lib/lru.ml:20:27:lt  s <= stamp → s < stamp
      20 │ | Some (_, s) when s <= stamp -> acc

    2 tests ran this line and none failed:
      lru › behaves like a list of bindings  test/test_cache.ml:44
      lru › evicts the least recently used   test/test_cache.ml:32
──────────────────────────────────────────────────────────

reproduce: dune exec --instrument-with ppx_windtrap.mutate test/test_cache.exe -- --arm lib/lru.ml:20:27:lt
mutants: 1 survived of 10 reached by this suite, 9 killed

In a project with several suites, dune exec windtrap -- mutants merges their verdicts into one report, the same way dune exec windtrap -- coverage does for coverage reports.

A survivor is either a missing test or a change that cannot matter. This one cannot matter: the eviction compares timestamps, and no two entries share one. You can annotate your code so that the line doesn't get mutated and the survivor won't show during the next run:

| Some (_, s) when (s <= stamp) [@mutate off "stamps are unique"] -> acc

Windtrap in practice

We found Windtrap invaluable in Invariant's projects and in Raven. The agents instinctively reach for the right kind of test (property tests, or stateful tests when possible and relevant), and we've written an agent skill that ships with Windtrap to guide them towards the test with the strongest oracle whenever they can, such as a stateful test instead of a list of unit tests.

Working on a project full of property tests and stateful tests, the agents write those kinds of tests themselves without being asked to, and materialize our intention to "multiply the good stuff".

Hunting bugs in the standard library

Still, we were curious to test our hypothesis that pushing our agents towards generative tests, such as property tests and stateful tests, leads to safer code. We're building on decades of research in the field, and on mature projects like QCheck and qcheck-stm, which were used to uncover bugs in the OCaml 5 runtime, so we had little doubt about the answer, but we still wanted to run the experiment.

We asked an agent to test the OCaml 5.5.1 standard library with Windtrap, and published the result at invariant-hq/windtrap-hunts.

Each technique found its own kind of bug:

Among the findings are also a few bugs of higher severity, which may have a security impact on OCaml programs. They are not included in the repository: we will disclose them privately to the OCaml security team, and will publish them once they are fixed.

We think the outcome of this experiment confirms the usefulness of Windtrap in moving towards safer, more trustworthy code written by agents. If agents equipped with Windtrap can find bugs, some of them serious, in the OCaml standard library, likely the most carefully reviewed and tested software in the OCaml ecosystem, it's a good tool for them to test their own code with.

That said, we also care very much about the user experience for humans, and put a lot of effort into making Windtrap's API and output as ergonomic and delightful as possible. And so we hope Windtrap will turn out to be a great tool for both your agents and yourself!

Windtrap is available on opam, with its PPX for expect tests, coverage and mutation testing:

opam install windtrap ppx_windtrap

Try it and don't hesitate to leave feedback at https://github.com/invariant-hq/windtrap/issues.

Happy hacking!