✨ Live browser view · session video · any model

Describe the test.
A real browser runs it.

ZeusQA turns a URL and a plain-English test into a real run. A frontier model drives Chromium step by step, checks what actually happened on screen, and hands you a verdict with checks, findings, screenshots and video.

No credit card. 5 free tests. Have an invite code?

zeusqa · live · llmgateway.ioLIVE
A finished ZeusQA run: verdict, checks, findings and a browser step with its screenshot
GPT-6 AstraClaude Sonnet 4.6Gemini 3 ProGPT-5.6 LunaDeepSeek V4Claude Haiku 4.5GPT-5.6 SolGemini 3.5 Flash270+ models
GPT-6 AstraClaude Sonnet 4.6Gemini 3 ProGPT-5.6 LunaDeepSeek V4Claude Haiku 4.5GPT-5.6 SolGemini 3.5 Flash270+ models
GPT-6 AstraClaude Sonnet 4.6Gemini 3 ProGPT-5.6 LunaDeepSeek V4Claude Haiku 4.5GPT-5.6 SolGemini 3.5 Flash270+ models
GPT-6 AstraClaude Sonnet 4.6Gemini 3 ProGPT-5.6 LunaDeepSeek V4Claude Haiku 4.5GPT-5.6 SolGemini 3.5 Flash270+ models

The problem

Scripts click. They don't look.

A recorded test passes while the page is broken, because it never checked what a person would see.

Selectors rot

Rename a button and a hundred scripts go red. A model reads the screen the way your users do and finds it anyway.

Console errors hide

The UI looks fine while the console fills with exceptions and 404s. Every run collects them as findings.

Nobody reads the artifacts

Screenshots pile up in CI. Here the model looks at each one before its next step, and the report shows you the ones that matter.

Everything ZeusQA does

One prompt, real evidence

Each capability, highlighted.

Prompt-driven testing

Write the test in plain English. No selectors, no framework. The model plans the steps and adapts when the UI changes.

Any model

GPT-6 Astra, Claude, Gemini or 270+ models through one gateway. Pick speed or depth per run.

Live browser view

Watch the real Chromium as the model drives it, at laptop, tablet or phone size. Stop a run at any point.

Session video + evidence

Every run records a video and a verdict: named checks, severity-rated findings, the exact scripts and the screenshots that back them.

Findings a script would miss

Console errors, failed requests and copy that contradicts behaviour — surfaced as findings, not buried in logs.

Organizations & projects

Group runs by project, share them across your team, and manage roles and billing per organization.

How it works

Three steps, one report

  1. 1

    Write the test in English

    A URL and what you want checked. Validation rules, a happy path, an edge case.

  2. 2

    Watch the browser work

    The model looks first, acts in small scripts, then looks again. Follow every step live.

  3. 3

    Read the verdict

    Pass, fail or blocked, with named checks, severity-rated findings and the evidence behind each one.

Pricing

Start free. Scale when you ship.

Every plan runs real browsers driven by frontier models. No credit card to start.

Free

Kick the tires.

$0/mo
Start free
  • 5 test runs (total)
  • 1 project
  • Live browser view
  • Community support

Pro

Most popular

For builders shipping fast.

$39/mo
Upgrade to Pro
  • 300 test runs / month
  • Unlimited projects
  • Live view + session video
  • Priority models (Astra)
  • Email support

Team

For teams that test together.

$149/mo
Upgrade to Team
  • 1,500 test runs / month
  • Organizations with members
  • Roles & shared projects
  • Everything in Pro
  • Priority support

Prices in USD. During the beta, checkout runs in Stripe test mode.

A real run

Sign-up flow on llmgateway.io, tested by Luna

PASS
verdict
46s
wall time
8/8
checks

Six scripts: it found the Get Started link, submitted the empty form, tried a bad email and a short password, created a throwaway account, then checked where the session landed and whether the console was clean.

info

The finding a script would have missed

The toast said check your email to verify before signing in, yet the session was redirected to onboarding with a working API key. Verification isn't enforced and the copy contradicts the behaviour — with the screenshot to prove it.

FAQ

Questions, answered

What is ZeusQA?
ZeusQA is prompt-driven QA for web apps. You give it a URL and a plain-English test; a frontier model drives a real Chromium browser, checks what happened on screen, and returns a verdict with checks, findings, screenshots and a session video.
How is it different from Playwright or Cypress?
Those run scripts you write and maintain. ZeusQA writes and adapts the steps itself from your description, reads the screen like a user, and reports console errors, failed requests and UI problems a fixed script would miss.
Which AI models can I use?
Any tool-calling model with vision on LLM Gateway — GPT-6 Astra, Claude, Gemini and 270+ others. Pick speed (Luna, Haiku) or depth (Astra) per run.
How much does it cost?
Free includes 5 test runs. Pro is $39/month for 300 runs, unlimited projects, live view and video. Team is $149/month for 1,500 runs and organizations with members.
Do I need to verify my email?
You can sign up and run free tests immediately. Email verification is required before subscribing to a paid plan.
Is it safe to run on my site?
The model treats page content as untrusted and avoids destructive actions unless your prompt allows them. It drives a real browser, so use staging or throwaway accounts for anything that writes data.

Your first test takes about a minute

Paste a URL, describe the flow, watch the browser do the rest.