An AI-agent-driven test platform · Free-style testing
An agent that tests your real code,
one round at a time.
Touchstone drives an AI agent through your repository. Point it at a project and a working directory, and it runs multi-round sessions that write test cases into a case library, execute them, and cluster failures into bug reports — all dispatched from one development board. Watchable, steerable, interruptible.
- 6task types
- 7stages
- 5board columns
- 1round = 1 session
- 1queue per project
- 0status polling
“On a concurrent duplicate submit, should we reject it or treat it as idempotent?”
This is the board you actually work in: one card is one dsh session, and test tasks and cards share a single queue — one thing runs per project at a time, and a card only gets a session when its turn comes. 5 columns · 1 queue
Not a sticky-note wall — a control surface
The development board lays out five columns: who is running, who is waiting, who is stuck, who is finished. It shares one queue with the test tasks — a project runs exactly one thing at a time.
One card is one session
A card in In progress means its session is running; a card in Blocked means the agent stopped to ask you something; a card in In review means the session went idle and the work is waiting for your call. Open the card and you get the whole session: thinking, tool calls, results — steerable mid-round, compactable, forkable from any round.
Cards and tasks share one queue
Test tasks, board cards and chat messages line up in a single queue, and the top-to-bottom order of the In-progress column is that queue. A card that hasn't come up yet shows Queued and only gets a session when its turn arrives. In a hurry, ⚡ Force jumps the queue at your own risk.
The agent stops and asks
Questions and approvals both go through one answer queue: the session suspends, the card lands in Blocked, and it resumes once the answer is delivered. Feishu pushes a card to your phone — reply “answer 390 1” or “approve 390” and it's answered. While it waits, the run slot is handed back and the head of the queue starts.
No session yet. Start puts the card in the queue, or pick 🌿 Start in a new worktree from the split button — that card runs immediately in its own git worktree: not queued, not holding the run slot, main checkout untouched.
The top half is the run zone, the bottom half the waiting area, split by a Waiting area · in queue order line. A running card takes a mid-round injection; a queued card can be stopped, which cancels the queue entry and drops it into review.
The agent asked a question or requested approval, and the session sits here until you answer. Answer or approve and it resumes by itself; nothing times out — the 60-second tick only reconciles state against reality.
The session is idle and the diff plus its conclusions are waiting for your judgement. Pass moves it to Done; Send back returns it with a reason — that is not an undo, it puts the card back on the development side to keep working.
A card moved here takes its dsh session into the archive with it; archive the session in dsh and the card walks here on its own. Reopen un-archives the session and sends the card back to review.
-
There are gates between the columns
Dragging a card into In progress queues it; an unfinished parent card blocks that and asks you to confirm Force start (at your own risk); a dependency cycle is refused outright. A project runs either Serial: one task at a time or Parallel: up to 5 at once.
-
Leaving the column leaves the queue
There is no separate "release the slot" action: drag a card away, stop it or delete it, and its queue row is finalised on the spot — the run slot goes to the next unit immediately.
-
Delete means recycle bin
Deleted cards go to the recycle bin and can be restored to their original column and position; only permanent deletion from the bin is irreversible. Emptying the bin is one click, and it takes the comments with it.
-
Paste images into a card, or pull cards from Jira
Card descriptions take pasted images and files (10 MB each, served behind auth). With Jira credentials configured, Import my to-dos turns the 50 issues assigned to you and still open into cards, de-duplicated by key.
-
The platform never commits for you
A card stops when the work is done; committing is yours, or the optional autocommit guard asset's. Every session on the board also carries its own permission preset: confirm each step, approve automatically, or fully autonomous.
Six task types, six kinds of output
Each type leaves something different behind: cases, a report, or just a record of one load run. The left column is the task type; the right column is what lands on disk.
Starts on a codebase it has never seen: map the modules, grade every feature P0/P1/P2, then generate de-duplicated cases along that order, run them one by one, and cluster the failures into bug reports by root cause. What to test is the agent's call, and coverage is written back after every round.
Writes no new cases, re-runs the ones you already have. Narrow it to what a commit date range touched, or re-run the whole library. Every case's verdict is written back to its own status file.
One agent round writes the plan: scope, target endpoints, load model, metric definitions, chart design and the command to reproduce it. The platform then drives the load itself — run it five times and you get five reports, each with its own logs and metrics.
Starts at report analysis: find the root cause, propose a fix, change the code, and redeploy and retest when the project says how. A change that was never built and deployed is not claimed as verified — the status stays honest and says so.
Pick an existing bug report and one of three scopes: retest only, redeploy then retest, or deploy only. When every related case already has a frozen verify.py, the platform runs the script instead and answers in seconds.
When you decide a report doesn't hold, the agent goes back to the related cases and corrects them: wrong assertions get fixed, cases that were false positives get deleted. The lesson goes into PITFALLS.md, which is read before any new case is written.
One round is one session
A task isn't "run a script". Each round hands work to a resident session that lives inside the dsh host process — one follow-up call, then wait for that round to finish.
One project runs one thing at a time
Tasks, board cards and chat messages for the same project share one queue, and whoever is running holds the slot. Different projects run in parallel, up to six. When a session is waiting on your answer it gives the slot up so the next item can start.
The session stays resident
There is no local agent subprocess and no status polling: one round is one follow-up plus one SSE turn/end. The session lives in the dsh host, and its id is recorded the moment the round starts.
You can cut in mid-run
Send queues a message for the next round; Inject steers it into the current one at the nearest step boundary. The session window shows every thought, tool call and result as it happens.
Seven stages
Every task picks a start and an end on the same track. The order on that track is a real order — this is a pipeline, not seven labels.
Solid spans are the default range; faded spans are optional end points. When Explore or Regression ends past "Report", the platform appends a follow-on task that carries on through analysis, fix, redeploy and retest — you don't create it yourself.
What's on disk after a round
The output is plain files, not rows locked inside a database. Read them, diff them, commit them — and the derived indexes can be deleted and rebuilt at any time.
- free_style/ # the case library, accumulated across rounds and sessions
- INDEX.md # feature grades · id cursor · totals
- ROUNDS.md # the angle and outcome of every round
- PITFALLS.md # false-positive lessons, read before writing cases
- login/session/
- FS0007_duplicate-login/
- case.md # the case definition (single source of truth)
- status.md # pass / fail / skipped / needs retest
- execution_20261007_0930.md # this run's request and response
- FS0007_duplicate-login/
- bug_report/
- 20261007_0930_FS_duplicate-login-still-passes/
- bug_report.md # symptom · evidence · root cause · related cases
- 20261007_0930_FS_duplicate-login-still-passes/
- .live/live.json # the monitor's data source (stage · progress · grades)
- .web/ # task logs · load-test metrics · board session logs
- board_media/ # images and files attached to board cards
-
De-duplication reads the library, not the agent's memory
A new case is checked against the library before it is accepted, and its id comes from the cursor in the root
INDEX.md. Only the main agent hands out ids, so two parallel lines can't take the same number. -
Derived indexes rebuild on demand
.live/cases.jsonl,similar.jsonlandrag_index.jsonlare all derived from the case library. Deleting them touches no source data. -
Monitoring and notifications ride on the same files
The monitor reads
live.jsonfor stage and progress, board cards carry their session, and Feishu sends a card when a task blocks or needs an answer — you can answer it by tapping on your phone. -
The library is searchable by meaning
With an embedding model configured, a retest ranks cases by semantic similarity to the changed files first. Without one it falls back to keyword matching, and never blocks a task over it.
What we don't do, and what we only half do
Half of a test platform's value is in the boundaries it will admit to. Here is what we deliberately don't do, or only do partly.
A fix stops at the file. If you want automatic commits, install the optional guard asset: by default it blocks once and tells the agent to commit, and only commits by itself when commitOnApproval is on.
kimi / claude / opencode / hermes and the bare dsh CLI are gone. An old binding fails loudly at start-up and tells you to rebind — it never silently degrades into a different runner.
The standalone site and the dsh plugin both default to ~/.touchstone/touchstone.db, guarded by a file lock so only one runs at a time. The second instance exits and prints who is holding it.
Both fork a new session from a given round; the old one is still there and you can switch back to it. They mean "branch from here", not "undo in place".
dsh has no event channel for it, so the session window shows no error banner. To see what didn't work, read the tool calls and results that actually happened.
The platform itself behaves the same on Windows and Linux — launcher, port fallback, path handling — but parent-death detection there has only the stdin pipe channel, and that has not been tested on a real machine.
Get it running
No environment yet? The live demo needs neither an install nor any data: create a project, create a task, watch it queue, answer the agent's question, and see cases and bug reports grow — all in the browser (a front-end simulation, no backend).
To run the real thing you need Python 3 and Node. Pick one of the two modes.
Standalone site
ships its own front end · 127.0.0.1:4601 by default
./touchstone.sh build # build the front end ./touchstone.sh start # start in the background ./touchstone.sh status # Windows: python touchstone.py start
Without webui/dist, server.py exits immediately — build first. The first start seeds one admin account: a random initial password is printed once in the start-up banner and must be changed at first login. The database lives at ~/.touchstone/touchstone.db and must be on a local disk — SQLite file locks don't work on network shares.
dsh plugin
embedded in dsh web · this mode or the site, not both
./dsh-plugin/install.sh # idempotent install into ~/.dsh/profiles/web # restart dsh web to take effect, then open the panel
This mode puts the panel in dsh's side rail. The agent session lives inside the host process, and the board, session window and live monitoring are identical to the standalone site. Only one of the two modes runs at a time.