Find what nothing asserts.
Across contract, code, cases,
autotests and CI.
A test management system with built-in agents that audits coverage across contract, code, cases, autotests, and CI runs. Kvalli adapts to your toolchain: connect multi-protocol contracts (OpenAPI, GraphQL SDL, gRPC Protobuf, Postman, HTTP files), scan backend code, draft test cases and native autotests steered by versioned Skills-as-Code, and autonomously heal broken locators or triage CI failures into your tracker.
--- name: "Idempotent Mutation Assertions" applies_to: ["testcase-agent", "codegen-agent"] triggers: tags: ["billing", "payments"] endpoints: ["/api/v1/orders", "/api/v1/checkout"] priority: 95 --- 1. Always assert 'Idempotency-Key' header on POST/PUT mutations. 2. Include a duplicate request test asserting 409 Conflict or cached replay.
A skill is a set of Markdown rules kept per project and edited in Kvalli. Agents automatically inject matching skills by domain triggers and priority to enforce your team conventions.
The workspace is live and open — no signup wall, no demo request. Anything on this page that is not built yet is marked as such in the status section. Start at /app.
- Stage
- MVP — liveworkspace open at /app, no signup wall
- Deployment
- Self-hosted single-tenantone organization per installation, telemetry off unless enabled
- Operation
- On demandagents run when you start them, on tasks, specs & CI results
- Trackers & Specs
- Jira · OpenAPI · GraphQL · gRPC · Postman · HTTPstories, multi-protocol schemas & backend code
- Code scan
- Express · Nest · FastAPI · Springextracts routes statically from source code
- Traceability
- 5 Dimensionscontract ↔ code ↔ cases ↔ autotests ↔ CI
- Self-Healing
- Autonomous Locator RecoveryDOM snapshot diffs & 1-click patches for Playwright/Cypress/Selenium
- Agents & Studio
- 4 Built-in + Custom Studiocustom agents per project, with dry-run gates
- Lock-in
- Lowautotests in your git, cases exportable to CSV/JSON
- Pricing
- Per project / reponot per-seat; pilot is free
- Test management
- Built-in, PostgreSQLimport from Qase, Allure TestOps
- Model engine
- BYOK (Gemini · OpenAI · Claude · Ollama)any OpenAI-compatible endpoint; deterministic fallback with no key
- Automation
- Playwright · Selenium across TS, Python, Java, Go; Cypress in TSTS/JS compiler-checked before the merge request
One feature, from task specification to filed bug
Agents run on demand across your delivery lifecycle: start one on a task or spec and it drafts test cases; start one on drafted cases and it writes native autotests; start one on a failed CI run and it triages it. Every artifact arrives as a merge request or a draft in Kvalli you review before it lands.
- You do
- Connect your task specification source (Jira task, Confluence PRD, or OpenAPI schema) and repository. Start the 5-dimension audit whenever you need it; it runs deterministically.
- It returns
- A unified matrix comparing specs to code: catches that POST /api/v2/orders/{id}/cancel is described in task PAY-231 and implemented in backend code (orders.controller.ts:84), but has 0 cases in your TMS, 0 autotests in your suite, and zero CI runs.
- Lands in
- A gap report in the coverage matrix before code is merged.
COVERAGE AUDIT · 5 DIMENSIONS VERIFIED
route ......... POST /api/v2/orders/{id}/cancel
spec source ... ✓ PAY-231 (Jira task / Confluence doc / OpenAPI)
backend code .. ✓ Implemented (orders.controller.ts:84)
tms cases ..... ✗ 0 cases matched in the test repository
autotests ..... ✗ 0 assertions in test suite
ci runs ....... ✗ Never executed in CI
GAP DETECTED
code .......... GAP-UNASSERTED-POST-_api_v2_orders__id__cancel
exposure ...... Critical business endpoint exists but lacks automated test assertions
recommend ..... Generate test cases from PAY-231 spec, emit autotests, and run in CI- You do
- Start the agent on task PAY-231 from Jira / Confluence. Your billing-domain skill is injected automatically based on domain tags.
- It returns
- Cases in your house style: Given–When–Then steps, your tag vocabulary, your priority rules, and an assertion on the error code envelope.
- Lands in
- A draft case in the built-in test repository, in review state, with its revision history.
---
id: "TC-205"
title: "POST /api/v2/orders/{id}/cancel — conflict on repeated cancel"
suite: "Orders & Checkout"
feature: "Cancellation"
priority: "high"
author: "Agent 1"
targetType: "api"
endpoint: "/api/v2/orders/{id}/cancel"
httpMethod: "POST"
sourceRef: "PAY-231 (Jira Story / Spec)"
appliedSkills: ["api-testcase-style", "billing-domain"]
---
# POST /api/v2/orders/{id}/cancel — conflict on repeated cancel
## Description
Derived from task specification in PAY-231: asserts idempotent handling when attempting to cancel an already-cancelled order.
## Steps to Reproduce
- 1. POST /api/v2/orders/ORD-991/cancel with a valid bearer token.
- 2. Verify the response status code is 409 Conflict.
- 3. Assert error response body matches { code: 'ORDER_ALREADY_CANCELLED' }.
## Expected Result
Returns 409 Conflict with structured error payload without triggering duplicate refunds.- You do
- Started for the test cases you select. Framework, language, and fixtures come from your repository configuration, not generic prompts.
- It returns
- A spec file that compiles directly into your suite: existing auth fixture reused, your import paths, and your describe naming.
- Lands in
- tests/api/orders/cancel.spec.ts (or your framework spec), committed into your git branch and opened as a merge request.
import { test, expect } from '@playwright/test';
import { createCancelledOrderFixture } from '../fixtures/orders.fixture';
test.describe('Orders API — cancellation lifecycle', () => {
test('returns 409 on duplicate cancel (PAY-231 / TC-205)', async ({ request }) => {
const { orderId, authToken } = await createCancelledOrderFixture(request);
const response = await request.post(`/api/v2/orders/${orderId}/cancel`, {
headers: { 'Authorization': `Bearer ${authToken}` }
});
expect(response.status()).toBe(409);
expect((await response.json()).error.code).toBe('ORDER_ALREADY_CANCELLED');
});
});- You do
- Your CI job uploads results with the Kvalli reporter; start triage on the failed run, and its logs and captured DOM snapshots are analysed.
- It returns
- A dual verdict: for backend regressions, drafts reproducible tickets to Jira/YouTrack. For frontend UI drift, analyzes the failure DOM snapshot, upgrades fragile selectors to Tier 1 test-ids, and emits unified diff patches.
- Lands in
- A bug ticket in your tracker, and a ready-to-merge self-healing code patch in your autotest branch.
CI TRIAGE & LOCATOR SELF-HEALING
verdict ....... UI text drift & locator timeout
broken ........ page.getByRole("button", { name: "Submit Order" })
dom match ..... <button data-testid="complete-order-btn">Complete Order</button>
RESILIENCE HEALING (96% CONFIDENCE · TIER 1)
--- tests/checkout.spec.ts (Original)
+++ tests/checkout.spec.ts (Healed)
@@ -14,1 +14,1 @@
- this.submitBtn = page.getByRole("button", { name: "Submit Order" });
+ this.submitBtn = page.getByTestId("complete-order-btn");
DEFECT DISPATCH
tracker ....... Jira / YouTrack (Auto-linked to PAY-231)
status ........ 1-click self-healing patch ready for branch PRIllustrative artifacts. The shape and field names match what the workspace emits; the values come from a sample project.
Built for a QA lead who cannot answer “what is not covered?”
The person this is aimed at owns a test suite they did not fully write, inherits conventions nobody documented, and gets asked in planning which parts of the API are actually tested. Below is the honest fit test, including where it fails.
Express, NestJS, FastAPI, Spring Boot, or an OpenAPI/Swagger spec. Kvalli extracts routes statically from source code or OpenAPI schemas — even if your documentation is outdated or absent (which surfaces shadow APIs).
Qase or Allure TestOps, with code on GitHub or GitLab and stories in Jira or YouTrack. Cases move over through importers with a preview; autotests stay in your repository.
Large enough that conventions matter and nobody holds the whole suite in their head; small enough that a QA lead can still say what the rules are.
That is the problem skills solve. If your team has no shared conventions yet, write them first — the tool enforces rules, it does not invent them.
Every artifact is a draft for review. If nobody has time to review, this adds work instead of removing it.
Code generation can still draft autotest specs (Playwright, Cypress, Selenium), but the coverage audit requires either a backend codebase (Express, Nest, FastAPI, Spring) or an OpenAPI contract to discover endpoints.
Kvalli is self-hosted, one organization per installation, by design. There is no multi-tenant SaaS and none is planned.
Wrong tool. Cases are structured steps, not Gherkin, and single sign-on is OIDC only; neither is on the roadmap.
Why this is not Copilot, Schemathesis, or an incumbent TMS
We win exactly two things: requirement-to-CI traceability across your existing stack, and versioned Skills-as-Code. Blindspot detection on OpenAPI contracts alone is a draw with free OSS tools. Below is the honest comparison across every tier.
Built-in, with revision history; export to CSV or JSON at any time.
Native test specs (Playwright, Cypress, Selenium) in your git branch.
Versioned Markdown skills per project, readable and editable in Kvalli.
| Axis | Free Contract OSS Schemathesis / Keploy | Incumbent TMS TestRail / Qase AI | Generic AI IDE Cursor / Copilot | Kvalli Traceability + Skills-as-Code |
|---|---|---|---|---|
| Blindspot detection | ⚠️ In contract boundary only (unaware of Jira requirements or TMS suites) | ❌ None (static manual repository) | ❌ None (only sees open files in IDE) | Correlates contract, code, TMS cases, autotests, and CI runs |
| Requirement linkage | ❌ None | ⚠️ Manual ticket link only | ❌ None | Full traceability: Jira/Confluence ↔ Contract ↔ Code ↔ CI result |
| Team conventions | ❌ Not applicable | ❌ None (generic text prompts) | ⚠️ Basic .cursorrules in prompt | Versioned per-project skills (Markdown rules) with trigger routing |
| Autotest codegen | ❌ Raw property-based tests only | ❌ None or vendor runner | ✅ Ad-hoc snippet in active file | Native tests (Playwright, Cypress, Selenium) in your git branch and review PR |
| TMS Integration | ❌ None | ⚠️ Isolated to vendor’s own database | ❌ None | Built-in test management with revision history, with lossless import from Qase and Allure TestOps |
| CI Failure triage | ❌ None | ⚠️ Basic log grouping | ❌ None | Automated triage: flake vs regression + live Jira/YouTrack ticket dispatch |
| Autotest Self-Healing | ❌ None (broken locators fail suite) | ❌ None | ❌ None (blind to runtime DOM snapshots) | Autonomous locator recovery: DOM snapshot diffs, Tier 1 test-id upgrades & 1-click patches |
| Model & Data perimeter | ✅ 100% local, no LLM required | ❌ Vendor cloud only | ❌ Vendor cloud proxy | BYOK (Gemini, OpenAI, Claude, any OpenAI-compatible endpoint, or local Ollama), deterministic fallback with 0 tokens |
Keep Schemathesis — it is great at raw RFC syntax compliance. But it has zero awareness of Jira user stories, existing TMS test suites, or shadow endpoints in repo code. Kvalli connects the full loop from business requirement to CI, ensuring you catch unasserted behavior outside the contract.
Vendor AI is locked inside their web UI: it cannot scan your backend git repository, cannot open autotest PRs in your branch, and will never integrate with competing trackers. Kvalli works across whichever tools your team already uses.
Security tools monitor production runtime traffic after deployment. Kvalli performs static route analysis of your source code, flagging undocumented endpoints and blind spots directly to QA before code ever reaches staging or production.
Connect your tools once. One step required to start, six optional.
Bring existing cases over with an importer that shows a preview before anything is written. Connect your code repository to run your first coverage audit, add tracker connections as needed, and write your team conventions down as versioned Markdown skills.
Where do specs, schemas, and PRDs live?
- LIVEOpenAPI / Swaggerv2 and v3
- LIVEGraphQL SchemasSDL query & mutation
- LIVEgRPC Protocol Buffers.proto RPC definitions
- LIVEPostman & HTTP filesv2.1 collections & RFC-7230
- LIVEConfluence PRDsCloud and Server/DC
- LIVEDirect route discovery from codeworks with zero specs
Requirements and schemas are pulled into normalized graph nodes (optional if starting from backend code).
Where is your product and test code hosted?
- LIVEGitHubpublic and private PAT
- LIVEGitLabGitLab.com and self-hosted
Scans implemented routes (Express/Nest/FastAPI/Spring) and autotest files over Git APIs.
Where do user stories and defects live?
- LIVEJiraREST v3 and v2
- LIVEYouTrack
- PARTIALLinearstories only, no defect filing
Optional. Pulls acceptance criteria into test cases and dispatches triaged bugs directly to your tracker.
Which system of record holds your test cases?
- LIVEBuilt-in repositoryversioned in PostgreSQL, no setup needed
- LIVEImport from Qase · Allure TestOpsTestOps read over its API; previewed before writing
Generated cases land in the built-in repository, with a revision kept for every change.
How should cases be stored and versioned?
- LIVEBuilt-in repositoryevery edit kept as a revision you can restore
- LIVEAutotests in your git repositoryreviewed by pull request
Cases keep their full history; the autotests generated from them go to your repository as merge requests.
What do your autotests actually run on?
- LIVEPlaywright · Cypress · Selenium
- LIVETypeScript · Python · Java · Go
- LIVEPage objects or direct scripts
Generated code compiles into your existing suite, not into a parallel runtime.
What are your team’s conventions and custom tool workflows?
- LIVEProject skills (Markdown rules)versioned per project; built-ins ship read-only, projects can override
- LIVEapplies_to · scope · priority · triggersper-agent & endpoint routing
- LIVECustom agents per projectwith dry-run gates
Every generation is steered by your conventions and custom tool integrations.
Cross-references task specifications (Jira, Confluence, OpenAPI) with routes extracted from backend code to find what is unasserted before merge.
Converts acceptance criteria into structured test cases obeying your versioned project skills.
Emits native autotest specs (Playwright, Cypress, Selenium) into a git branch, reusing your auth fixtures and page objects.
Ingests CI run logs, diagnoses failure root cause, classifies flakiness vs regressions, drafts defect tickets, and autonomously heals broken UI locators across Playwright, Cypress, and Selenium via DOM snapshot diffs.
These are not autonomous hires — they are drafting pipelines. Every artifact (a test case, an autotest spec in your framework, a ticket) arrives as a pull request in your repository or a draft in Kvalli for your team to review.
What is built, what is half-built, what is not built
Some of this is still a stub, and we say which. The matrix below is the honest state of the repository, and it is the same list we work from.
Pluggable interfaces over Jira, YouTrack, Confluence, GitHub, and GitLab so the tool joins your stack.
Deterministic evaluation across contracts, product code, TMS cases, autotests, and CI runs.
Bring-your-own-key routing across four vendors plus any OpenAI-compatible endpoint, Skills-as-Code prompt steering, and deterministic fallback.
Test cases, runs, skills and custom agents in PostgreSQL, with revision history; Git tree caching for repository scans.
| LIVE | Agent 0 — Coverage audit engine | 5D matrix across multi-format contracts (OpenAPI, GraphQL SDL, gRPC Protobuf, Postman, HTTP files), product code (Express/Nest/FastAPI/Spring), TMS cases, autotests, and CI runs. |
| LIVE | Agent 1 — Requirements & Spec-to-Cases | Pulls stories from Jira/YouTrack/Confluence and generates zero-hallucination test suites across OpenAPI, GraphQL, gRPC, and Postman specs. |
| PARTIAL | Agent 2 — Multi-framework autotest codegen | Page Object Model code generation for Playwright and Selenium in TypeScript, Python, Java and Go, and for Cypress in TypeScript (its only language). TS/JS is compiler-checked before the merge request; Python is parsed, and Java and Go are not checked. |
| LIVE | Agent 3 — CI failure triage & self-healing | On-demand failure triage, flakiness diagnosis, defect ticket dispatch, and autonomous DOM locator healing studio with unified diffs. |
| LIVE | Git provider repo sync | GitHub and GitLab REST API tree & file synchronization with local caching. |
| LIVE | Confluence adapter | Atlassian Confluence Cloud & Server REST API integration with storage HTML parsing. |
| LIVE | Multi-LLM router | Gemini, OpenAI, Claude, Ollama and any OpenAI-compatible endpoint (e.g. DeepSeek). |
| LIVE | Defect tracker dispatch | Live REST ticket creation for Jira (v3/v2) and YouTrack. |
| LIVE | Built-in test repository | Cases with review states and a stored version history, in PostgreSQL. |
| LIVE | Custom QA Agents Studio | Declarative agent configurations with closed output types, Skills-as-Code trigger routing, and dry-run gates. |
| LIVE | Security hardening | SSRF filtering with a trustedHosts whitelist, sliding-window rate limits, and integration credentials encrypted with AES-256-GCM in PostgreSQL; the API never returns them unmasked. |
| LIVE | Import from Allure TestOps | Cases, steps, shared steps, trees, owners and files read from the TestOps API, previewed before writing. |
| LIVE | CI pipeline trigger | Starts GitLab CI pipelines and GitHub Actions workflows directly from Kvalli. |
| LIVE | Auth, roles & OIDC SSO | Users, roles and single sign-on over OIDC. One organization per installation, by design. |
The harness seeds known specification violations into three open-source API contracts and runs generation twice — once bare, once with the repository's skills — then scores which seeded defects each arm asserted.
We are not publishing a synthetic recall number yet. The committed snapshot ran without an active model key, so both arms fell back to the deterministic schema parser and caught 0 of 15 seeded edge-case specification defects — in both arms. That measures the deterministic fallback, not the difference skills make.
npm test
Verifies cross-source resolution across contracts, code, cases, autotests, and CI runs, Skills-as-Code frontmatter parsing, trigger routing, custom agents dry-run verification gates, and parser fixtures.
Kvalli is distributed as an open-core, self-hosted instance. Pilot design partners receive private repository access, Docker images, and dedicated onboarding. Defect recall across the 15 seeded specification defects with a production model key has not been published yet; npm run eval:publish produces it, and the figures will appear here once it has run against a real model.
To evaluate on your own codebase with your own BYOK key, run the live evaluation harness in /app workspace.
What leaves your infrastructure, what is stored, and where code runs
For engineering teams granting read access to backend repositories and issue trackers: Kvalli is architected so your code, credentials, and test assets remain strictly inside your security boundary.
Kvalli runs as a single-tenant Node.js process or container on your infrastructure. All source parsing, route extraction, and test indexing execute locally.
You supply your own model key (Gemini, Claude, OpenAI or any OpenAI-compatible endpoint) or run offline with local Ollama. The deterministic fallback runs with 0 tokens and 0 bytes sent to external APIs.
Git, Jira and LLM API tokens are encrypted with AES-256-GCM in PostgreSQL under a versioned key. The API never returns unmasked credentials to clients or browsers.
Built-in SSRF protection blocks private RFC-1918 subnets and cloud instance metadata (169.254.169.254), with explicit trustedHosts whitelisting for self-hosted GitLab/Jira.
Three to five teams, running it on their own contracts
Pilot is free. Beyond pilot, pricing will be per project/repository rather than per seat — because quality governance is owned by leads and CI, not 40 seats. In exchange: direct onboarding, skills tuned to your conventions, and priority on your team's adapters.