Property Testing with Agent Swarms

Have agents build the machinery to find bugs in your repo, then let that machinery generate and check cases without spending tokens on each one. You can do this with ordinary coding agents, without access to a closed cybersecurity program. Writing reference implementations, generators, and assertions for every custom data structure and algorithm is now work you can hand to a swarm.

Pick a complex repo at work. Ask the people who know it well which parts worry them. Give a strong planning agent the property-testing skill from my agent-skills repo and those leads. It gives the agent methods for choosing targets, building reference models and generators, and turning failures into reviewable fixes. Target custom data structures, query planners, graph algorithms: complicated behavior with a simple way to check correctness.

[Read More]

Fresh Eyes as a Service: Using LLMs to Test CLI Ergonomics

Every tool has a complexity budget. Users will only devote so much time to understanding what a tool does and how before they give up. You want to use that budget on your innovative new features, not on arbitrary or idiosyncratic syntax choices.

Here’s the problem: you can’t look at it with fresh eyes and analyze it from that perspective. Finding users for a new CLI tool is hard, especially if the user experience isn’t polished. But to polish the user experience, you need users, you need user feedback, and not just from a small group of power users. Many tools fail to grow past this stage.

What you really need is a vast farm of test users that don’t retain memories between runs. Ideally, you would be able to tweak a dial and set their cognitive capacity, press a button and have them drop their short term memory. But you can’t have this, because of ‘The Geneva Convention’ and ’ethics’. Fine.

LLMs provide exactly this. Fresh eyes every time (just clear their context window), adjustable cognitive capacity (just switch models), infinite patience, and no feelings to hurt. Best of all, they approximate the statistically average hypothetical user - their training data includes millions of lines of humans interacting with a wide variety of CLI tools. Sure, they get confused sometimes, but that’s exactly what you want. That confusion is data, and the reasoning chain that led to it is available in the LLM’s context window.

[Read More]
llm  ux  cli  testing  detect  tools