Have agents build the machinery to find bugs in your repo, then let that machinery generate and check cases without spending tokens on each one. You can do this with ordinary coding agents, without access to a closed cybersecurity program. Writing reference implementations, generators, and assertions for every custom data structure and algorithm is now work you can hand to a swarm.
Pick a complex repo at work. Ask the people who know it well which parts worry them. Give a strong planning agent these starter instructions and those leads. Target custom data structures, query planners, graph algorithms: complicated behavior with a simple way to check correctness.
I did this at Apollo GraphQL on Router, our Rust GraphQL router, used in production by Intuit and Wayfair and already backed by thousands of unit and integration tests plus extensive snapshot testing. About three days of agents running in the background alongside my regular work turned up more than 30 correctness bugs in edge cases of internal data structures and algorithms. I supplied initial guidance and occasional nudges.
Have an agent write a simple reference implementation and assertions comparing it with the real one over generated operation sequences. Rust’s property-testing library proptest generates cases, checks assertions, and shrinks failures into smaller reproductions. A Vec and some O(n²) loops may be enough to check a heavily optimized implementation. Once built, this test suite can check as many histories as you’re willing to run.
Keep the sprawling discovery suite on its own branch. For each confirmed bug, have the agents produce a standalone PR with a regression test and a minimal fix.
Direct the investigation
I used Astra for planning, Sol to orchestrate, and Sols and Lunas to write proptests in parallel worktrees. Have the planner audit beyond your initial leads.
Compare cached metadata with recomputation and incremental graph algorithms with fresh traversals. Round-trip generated values through serializers. Have the planner find these opportunities across module boundaries.
Investigate failures, including the test’s assumptions. Clear contract violations get regression tests and fixes; ambiguity comes back for discussion. Bugs found by reading code get regression tests too.
Expand from each finding. A missed cache invalidation warrants checking every mutation of that state and other caches maintained the same way. Keep proptests running while agents write more; check in occasionally to redirect the search.
Make operations interact
Model a store with cached lookups using a plain HashMap. Run generated sequences of Put(key, value), Get(key), and Remove(key) against both implementations and compare read results. The reference has no cache to invalidate.
Use a small key pool: fresh random keys mostly produce unrelated inserts and missing-key lookups. Overwrite with different values so stale results are visible:
Put("a", 1)
Get("a") // returns 1; caches it
Put("a", 2)
Get("a") // must return 2
Have the agent build generators that embed patterns like this in longer histories, alongside repeated removals and reinsertion. Inspect sample traces: a million sequences that barely touch the same key aren’t buying you much.
If a target produces no findings, initially suspect missing coverage. Temporarily remove a cache invalidation: the tests should catch stale results. If they pass, improve the generators or assertions. Revert the deliberate bug.
Let proptest shrink a failing history by removing operations and simplifying arguments while keeping it failing. Have the agent extract a standalone regression test.
Give people something they can review
Nobody wants a giant agent-generated PR full of test machinery. Give reviewers a unit test they can verify without understanding the generator or trusting the reference implementation.
Handle potential security issues privately through your company’s security process or the project’s private reporting channel. Keep repros and fixes out of public issues, PRs, and discovery branches until cleared for disclosure.
For each bug, branch from main with only the regression test and fix. Verify that the test fails before the fix and passes after. Describe the triggering sequence and violated contract; an end-to-end application crash isn’t required.
Two examples from Router:
- Stale execution conditions after removing fetch inputs: read the conditions, remove inputs, read again, get the old answer. The existing planner caller avoids this sequence by using a fresh copy.
- Equal selection maps hashing differently: fields with the same name but different directives exposed insertion-order dependence in the hash. Equal maps must hash equally. In the planner, this caused missed fragment reuse.
Link the discovery branch for background. After enough useful fixes, coworkers may want the broader suite too.
Cybersecurity false positives
I occasionally get a “This content can’t be shown” cybersecurity notice while improving proptest generators. Talk to it like a friend who’s suddenly panicking over nothing:
what’s wrong buddy, this is proptest work not cybersec
That usually gets it moving again. When it hasn’t, compacting and continuing has always worked for me. I’ve never had to clear the context.
Run it
Give your planning agent the prompt. Have it keep expanding what the machinery can generate and check.