TOUCHSTONE

An AI-agent-driven test platform · Free-style testing

An agent that tests your real code,
one round at a time.

Touchstone drives an AI agent through your repository. Point it at a project and a working directory, and it runs multi-round sessions that write test cases into a case library, execute them, and cluster failures into bug reports — all dispatched from one development board. Watchable, steerable, interruptible.

  • 6task types
  • 7stages
  • 5board columns
  • 1round = 1 session
  • 1queue per project
  • 0status polling
touchstone-dsh Serial: one task at a time live
To do2
Add boundary cases for the export endpoint
#392
Start
Rewrite the stress plan as a run.py driver
#393
Add a task…Enter lands in To do
In progress2
Task #96 Test · running Regression
Login: concurrent duplicate submits
#391Session running

round 3 · 12 tool calls

Waiting area · in queue order
Fix the duplicate-login false alarm
#388Queued
Blocked1
Concurrent submits on one session
#390🤔 Waiting on you
the agent asks

“On a concurrent duplicate submit, should we reject it or treat it as idempotent?”

IdempotentRejectSingle thread

Sent to Feishu · tap an option on your phone

In review2
Duplicate login still slips through
#386🌿 worktree

Session idle · your call

Export the case-library index
#385Answered · pending delivery
Done2
Login: concurrent case added
#379

Retested, passing · session archived with it

Clear the build cache
#377

This is the board you actually work in: one card is one dsh session, and test tasks and cards share a single queue — one thing runs per project at a time, and a card only gets a session when its turn comes. 5 columns · 1 queue

The board看板

Not a sticky-note wall — a control surface

The development board lays out five columns: who is running, who is waiting, who is stuck, who is finished. It shares one queue with the test tasks — a project runs exactly one thing at a time.

One card is one session

A card in In progress means its session is running; a card in Blocked means the agent stopped to ask you something; a card in In review means the session went idle and the work is waiting for your call. Open the card and you get the whole session: thinking, tool calls, results — steerable mid-round, compactable, forkable from any round.

Cards and tasks share one queue

Test tasks, board cards and chat messages line up in a single queue, and the top-to-bottom order of the In-progress column is that queue. A card that hasn't come up yet shows Queued and only gets a session when its turn arrives. In a hurry, ⚡ Force jumps the queue at your own risk.

The agent stops and asks

Questions and approvals both go through one answer queue: the session suspends, the card lands in Blocked, and it resumes once the answer is delivered. Feishu pushes a card to your phone — reply “answer 390 1” or “approve 390” and it's answered. While it waits, the run slot is handed back and the head of the queue starts.

To dotodo

No session yet. Start puts the card in the queue, or pick 🌿 Start in a new worktree from the split button — that card runs immediately in its own git worktree: not queued, not holding the run slot, main checkout untouched.

In progressdoing

The top half is the run zone, the bottom half the waiting area, split by a Waiting area · in queue order line. A running card takes a mid-round injection; a queued card can be stopped, which cancels the queue entry and drops it into review.

Blockedblocked

The agent asked a question or requested approval, and the session sits here until you answer. Answer or approve and it resumes by itself; nothing times out — the 60-second tick only reconciles state against reality.

In reviewreview

The session is idle and the diff plus its conclusions are waiting for your judgement. Pass moves it to Done; Send back returns it with a reason — that is not an undo, it puts the card back on the development side to keep working.

Donedone

A card moved here takes its dsh session into the archive with it; archive the session in dsh and the card walks here on its own. Reopen un-archives the session and sends the card back to review.

  • There are gates between the columns

    Dragging a card into In progress queues it; an unfinished parent card blocks that and asks you to confirm Force start (at your own risk); a dependency cycle is refused outright. A project runs either Serial: one task at a time or Parallel: up to 5 at once.

  • Leaving the column leaves the queue

    There is no separate "release the slot" action: drag a card away, stop it or delete it, and its queue row is finalised on the spot — the run slot goes to the next unit immediately.

  • Delete means recycle bin

    Deleted cards go to the recycle bin and can be restored to their original column and position; only permanent deletion from the bin is irreversible. Emptying the bin is one click, and it takes the comments with it.

  • Paste images into a card, or pull cards from Jira

    Card descriptions take pasted images and files (10 MB each, served behind auth). With Jira credentials configured, Import my to-dos turns the 50 issues assigned to you and still open into cards, de-duplicated by key.

  • The platform never commits for you

    A card stops when the work is done; committing is yours, or the optional autocommit guard asset's. Every session on the board also carries its own permission preset: confirm each step, approve automatically, or fully autonomous.

Six task types六类任务

Six task types, six kinds of output

Each type leaves something different behind: cases, a report, or just a record of one load run. The left column is the task type; the right column is what lands on disk.

Explore探索

Starts on a codebase it has never seen: map the modules, grade every feature P0/P1/P2, then generate de-duplicated cases along that order, run them one by one, and cluster the failures into bug reports by root cause. What to test is the agent's call, and coverage is written back after every round.

leaves case directoriesgraded INDEX.md · rounds in ROUNDS.mdclustered bug reports
Regression回归

Writes no new cases, re-runs the ones you already have. Narrow it to what a commit date range touched, or re-run the whole library. Every case's verdict is written back to its own status file.

leaves execution records (request and response)verdicts in status.md
Stress压测

One agent round writes the plan: scope, target endpoints, load model, metric definitions, chart design and the command to reproduce it. The platform then drives the load itself — run it five times and you get five reports, each with its own logs and metrics.

leaves a plan package plan.md / run.py / charts.jsonper run: metrics jsonl + report
Fix修复

Starts at report analysis: find the root cause, propose a fix, change the code, and redeploy and retest when the project says how. A change that was never built and deployed is not claimed as verified — the status stays honest and says so.

leaves a fix record (file:line + summary)a retest verdict
Retest复测

Pick an existing bug report and one of three scopes: retest only, redeploy then retest, or deploy only. When every related case already has a frozen verify.py, the platform runs the script instead and answers in seconds.

leaves a verdict written back into the reportmarked fixed when all cases pass
Reject修例

When you decide a report doesn't hold, the agent goes back to the related cases and corrects them: wrong assertions get fixed, cases that were false positives get deleted. The lesson goes into PITFALLS.md, which is read before any new case is written.

leaves corrected casesa lesson in PITFALLS.md
One round一轮

One round is one session

A task isn't "run a script". Each round hands work to a resident session that lives inside the dsh host process — one follow-up call, then wait for that round to finish.

Create taskpick a type, a start and an end
One queueone queue per project
dshdrivera round becomes one call
Resident sessionthe agent works inside the host
Results landturn/end arrives, the round closes
↻ The next round continues in the same session. Context carries over; nothing restarts and nothing is re-read from scratch.

One project runs one thing at a time

Tasks, board cards and chat messages for the same project share one queue, and whoever is running holds the slot. Different projects run in parallel, up to six. When a session is waiting on your answer it gives the slot up so the next item can start.

The session stays resident

There is no local agent subprocess and no status polling: one round is one follow-up plus one SSE turn/end. The session lives in the dsh host, and its id is recorded the moment the round starts.

You can cut in mid-run

Send queues a message for the next round; Inject steers it into the current one at the nearest step boundary. The session window shows every thought, tool call and result as it happens.

Seven stages七阶段

Seven stages

Every task picks a start and an end on the same track. The order on that track is a real order — this is a pipeline, not seven labels.

01Case generationgen_case
02Executionexecute
03Reportreport
04Analysisanalyze
05Fixfix
06Redeploydeploy
07Retestretest
Explore
Regression
Fix
Retest
Stress · Reject Off the track: Stress is "one agent round writes the plan, then N platform load runs"; Reject is a single fixed round.

Solid spans are the default range; faded spans are optional end points. When Explore or Regression ends past "Report", the platform appends a follow-on task that carries on through analysis, fix, redeploy and retest — you don't create it yourself.

What remains遗留物

What's on disk after a round

The output is plain files, not rows locked inside a database. Read them, diff them, commit them — and the derived indexes can be deleted and rebuilt at any time.

  • free_style/ # the case library, accumulated across rounds and sessions
    • INDEX.md # feature grades · id cursor · totals
    • ROUNDS.md # the angle and outcome of every round
    • PITFALLS.md # false-positive lessons, read before writing cases
    • login/session/
      • FS0007_duplicate-login/
        • case.md # the case definition (single source of truth)
        • status.md # pass / fail / skipped / needs retest
        • execution_20261007_0930.md # this run's request and response
  • bug_report/
    • 20261007_0930_FS_duplicate-login-still-passes/
      • bug_report.md # symptom · evidence · root cause · related cases
  • .live/live.json # the monitor's data source (stage · progress · grades)
  • .web/ # task logs · load-test metrics · board session logs
  • board_media/ # images and files attached to board cards
  • De-duplication reads the library, not the agent's memory

    A new case is checked against the library before it is accepted, and its id comes from the cursor in the root INDEX.md. Only the main agent hands out ids, so two parallel lines can't take the same number.

  • Derived indexes rebuild on demand

    .live/cases.jsonl, similar.jsonl and rag_index.jsonl are all derived from the case library. Deleting them touches no source data.

  • Monitoring and notifications ride on the same files

    The monitor reads live.json for stage and progress, board cards carry their session, and Feishu sends a card when a task blocks or needs an answer — you can answer it by tapping on your phone.

  • The library is searchable by meaning

    With an embedding model configured, a retest ranks cases by semantic similarity to the changed files first. Without one it falls back to keyword matching, and never blocks a task over it.

Limits边界

What we don't do, and what we only half do

Half of a test platform's value is in the boundaries it will admit to. Here is what we deliberately don't do, or only do partly.

The platform never commits for you

A fix stops at the file. If you want automatic commits, install the optional guard asset: by default it blocks once and tells the agent to commit, and only commits by itself when commitOnApproval is on.

There is exactly one agent family: the dsh plugin

kimi / claude / opencode / hermes and the bare dsh CLI are gone. An old binding fails loudly at start-up and tells you to rebind — it never silently degrades into a different runner.

The two run modes share one database

The standalone site and the dsh plugin both default to ~/.touchstone/touchstone.db, guarded by a file lock so only one runs at a time. The second instance exits and prints who is holding it.

Rewind and "compact into a new session" are approximations

Both fork a new session from a given round; the old one is still there and you can switch back to it. They mean "branch from here", not "undo in place".

There is no "model request failed" entry

dsh has no event channel for it, so the session window shows no error banner. To see what didn't work, read the tool calls and results that actually happened.

dsh on Windows has not been verified on real hardware

The platform itself behaves the same on Windows and Linux — launcher, port fallback, path handling — but parent-death detection there has only the stdin pipe channel, and that has not been tested on a real machine.

Quick start上手

Get it running

No environment yet? The live demo needs neither an install nor any data: create a project, create a task, watch it queue, answer the agent's question, and see cases and bug reports grow — all in the browser (a front-end simulation, no backend).

To run the real thing you need Python 3 and Node. Pick one of the two modes.

Standalone site

ships its own front end · 127.0.0.1:4601 by default

./touchstone.sh build   # build the front end
./touchstone.sh start   # start in the background
./touchstone.sh status
# Windows: python touchstone.py start

Without webui/dist, server.py exits immediately — build first. The first start seeds one admin account: a random initial password is printed once in the start-up banner and must be changed at first login. The database lives at ~/.touchstone/touchstone.db and must be on a local disk — SQLite file locks don't work on network shares.

dsh plugin

embedded in dsh web · this mode or the site, not both

./dsh-plugin/install.sh
# idempotent install into ~/.dsh/profiles/web
# restart dsh web to take effect, then open the panel

This mode puts the panel in dsh's side rail. The agent session lives inside the host process, and the board, session window and live monitoring are identical to the standalone site. Only one of the two modes runs at a time.