Kallisti Contact

ΚΑΛΛΙΣΤΗΙ

We play your build all night and send you what broke.

Kallisti is an AI-agent studio. Our agents explore your browser game or web product without a pre-written test script, across desktop and mobile viewports, and record what they find. You get a written report with severity, evidence, and steps to reproduce.

  • 39 games built and swept in house
  • Desktop and mobile viewports
  • Recorded frames and repro steps
  • Sweeps from $149

I.  What we do

Three service lines. One of them is the reason to hire us.

Flagship

Agentic QA and regression playtesting

Autonomous agents play your build overnight. A scripted suite catches a failure once somebody writes the case for it. Our agents are not following your test plan, so they end up in the states nobody wrote a case for: dead ends, screens that stop responding to input, silent console errors, failed asset and network requests, layout that collapses at phone width.

Every finding carries a severity, a recorded frame, and the exact input sequence that produced it. Run the sweep again after your next release and the report tells you what changed across the states both runs reached.

The report separates confirmed failures from heuristic signals and states plainly what the run did not cover. You should be able to hand it to a developer without translating it first.

A sweep is a time-boxed black-box run in headless Chromium at a desktop viewport and an emulated mobile viewport. It reports the defects it observed inside that window. It does not prove there are no others, and it does not cover security, load, accessibility conformance, cross-browser parity, or anything server-side.

Game sweeps run $149 to $299 per sweep, depending on scope and title count. Web-app sweeps run $299 to $499: the same overnight pass under an explicit safety policy — staging targets by default, destructive actions blocked, metered endpoints budgeted — with an authenticated test account you provide. One sweep is one build, both viewport passes, and one written report.

In every sweep

  • Dead-end and non-response detection
  • Console and page-error capture
  • Failed network and asset requests
  • Desktop and emulated mobile passes
  • Recorded frames per finding
  • Exact reproduction sequences
  • Diff of visited states across versions
  • Stated coverage boundaries

Service line

Custom agent integrations

Narrow MCP connectors that let an AI assistant operate the tools a business already runs. Claude or Cursor reads and writes real records in the workflow platform instead of describing what someone should go do by hand.

First vertical is field service software, starting with Jobber. Built as a fixed engagement, then kept working on retainer.

Priced per platform. Talk to us.

Service line

Outcome services

Recurring research and data work, delivered as a finished result rather than billed hours. Prospect research, competitive intelligence, catalog cleanup.

You agree on the output and the schedule. The file arrives on the schedule.

Paid pilots from $250.

II.  How it works

You see the work before you pay for it.

This is the whole sales process. We do not pitch a capability. We run the sweep against your own product and send you the finished report, then you decide whether it was worth buying.

  1. 1 / 4

    You send a URL

    A build link, a staging environment, or a published game. No integration work, no code access, nothing to install.

  2. 2 / 4

    We sweep it overnight

    Agents drive the product in headless Chromium at a desktop viewport and an emulated mobile viewport, from seeded exploration runs. Every input is logged with the frames recorded beside it.

  3. 3 / 4

    You read the report

    Findings, severity, recorded evidence, repro steps, and the limits of the run. It arrives before you have paid anything.

  4. 4 / 4

    You decide

    If the report earned its place in your release process, we set a cadence and a price. If it did not, you keep the findings and we part on good terms.

III.  The artifact

Read a real report before you talk to us.

Below is an unedited report from a sweep of our own test corpus. It is the same document a client receives. It shows the format and the depth of the findings on our own games rather than an outside team’s verdict on them. Read the findings, then read the boundaries section, which says what the run did not check.

A clean result means no defect was observed in the recorded time box.

From the report’s boundaries section
Open the sample report

What is inside

  • 01An executive readout with counts by severity
  • 02Each finding classified as failure or heuristic signal
  • 03The exact timestamped input sequence to reproduce it
  • 04Recorded frames linked beside every finding
  • 05Console errors, page errors, and failed requests
  • 06A mobile geometry and overflow pass at 375 by 812
  • 07A persistence probe across reload
  • 08A methodology section and a stated list of boundaries

IV.  Who we are

One operator directing a fleet of AI agents.

A one-person studio is the reason the pricing works. The agents run the sweeps, the research, and the integration builds. The operator sets the scope, checks the output against the product, and signs the report. Nothing goes to a client that a person has not read.

We built the QA harness for our own games first, which is why the corpus is real and the report format survived contact with actual defects.

Test corpus
39 browser games built and QA’d in house.
Commissioned work
Shipped paid digital board game commissions for outside clients.
Method
Seeded, deterministic exploration. Every input sequence is recorded and replayable, though timing and network conditions can shift what a replay produces.

V.  Contact

Send a URL. We will send back a report.

Tell us what the product is and where it lives. Most browser games and web products can be swept as they are. If yours needs a login, a seeded account, a multiplayer session, or a payment step, say so and we will tell you before you send anything whether a sweep fits. For integration or research work, describe the tool and the outcome you want.

Before you send a link

Build links
We sweep only the URL you send. It is not published, not passed to a third party, and not used to train anything. The sweep drives a local browser on our own hardware, so your build is not uploaded to a model provider. Ask and every artifact from your sweep is deleted, frames included.
Non-destructive by default
Agents read and play. They do not create accounts, submit contact or payment forms, spend inventory, or generate load. Anything that writes to a live system happens only against a target you nominate first, in writing.
Who reads it
The operator reads every report before it ships, checks the findings against the build, and marks the heuristic ones that still need a human to confirm them.
Free to paid
The first sweep and its report are free. No invoice, no obligation, no card. You pay only if you decide to put sweeps into your release process, and then it is $149 to $299 per sweep.
support@kallisti.app

The sweep runs overnight. The report normally lands within two working days of the build link.