I Built a Tool to Find the Tests Developers Have Stopped Trusting
A test fails in CI.
You look at it.
You run it again.
It passes.
You run it again.
It passes.
You rerun the pipeline.
Everything is green.
So you merge the pull request.
A few hours later, the same test fails again.
At some point, the team stops treating that failure as useful information.
That’s the real problem with flaky tests.
They don’t just make CI unreliable.
They make developers stop trusting CI.
The problem isn’t writing tests
Modern JavaScript projects already have excellent testing tools.
Playwright is great for browser testing.
Vitest and Jest are great for unit and integration testing.
CI platforms can execute thousands of tests.
The problem starts after the tests exist.
A growing project eventually accumulates:
- Tests that fail intermittently
- Tests that take too long
- Tests that haven’t failed in months
- Tests that fail only in CI
- Tests that fail because of timing
- Tests that produce almost identical errors
- Tests that nobody wants to touch
The test suite becomes a second system that needs maintenance.
And most teams don’t have a good way to measure its health.
So I built TestOps Kit
The idea is simple:
Don’t replace the test runner. Analyze what it produces.
TestOps Kit consumes machine-readable test results and builds a reliability layer around them.
The workflow looks like this:
Playwright / Vitest / Jest
↓
Test results
↓
TestOps Kit
↓
┌──────────┼───────────┐
↓ ↓ ↓
History Reliability Performance
↓ ↓ ↓
Flaky Quarantine Slow tests
tests
↓
Dashboard
Enter fullscreen mode Exit fullscreen mode
The existing testing workflow stays intact.
TestOps simply adds another layer of information.
The first feature I wanted was flaky-test detection
A single failed test doesn’t necessarily mean a test is flaky.
You need history.
For example:
Run 1 → PASS
Run 2 → PASS
Run 3 → FAIL
Run 4 → PASS
Run 5 → FAIL
Enter fullscreen mode Exit fullscreen mode
That pattern is much more interesting than a single failure.
TestOps tracks test behavior across runs and calculates reliability metrics from the available history.
The goal isn’t to say:
“This test is definitely broken.”
The goal is to say:
“This test has demonstrated inconsistent behavior and deserves investigation.”
That’s a much more useful signal.
Then came quarantine
Once you identify unreliable tests, the next problem is operational.
What do you do with them?
TestOps can generate a quarantine manifest based on a configurable flakiness threshold.
For example:
Flaky tests
checkout/payment.spec.ts 42%
account/login.spec.ts 31%
orders/create.spec.ts 27%
Enter fullscreen mode Exit fullscreen mode
Instead of relying on someone’s memory, the team now has a concrete list.
The important part is that quarantine is treated as a temporary reliability workflow, not a way to permanently hide failures.
Slow tests are another form of test debt
A test doesn’t have to fail to become expensive.
Imagine a suite containing:
1,200 tests
Enter fullscreen mode Exit fullscreen mode
and a handful of tests account for a significant portion of the runtime.
Those tests affect every developer.
Every pull request.
Every CI run.
Every deployment.
So TestOps also tracks duration and highlights slow tests.
The goal is straightforward:
Find the tests that are costing the team time.
Visual regression belongs in the same conversation
Testing isn’t only about pass and fail.
UI changes can also introduce unexpected regressions.
TestOps includes deterministic snapshot checking so a CI pipeline can detect changed snapshot content.
For example:
npx testops snapshot snapshots --strict
Enter fullscreen mode Exit fullscreen mode
If a snapshot changes unexpectedly, the command returns a failure.
That makes it possible to use the same reliability workflow in CI.
I also wanted a boring CLI
Developer tools don’t need complicated installation procedures.
The basic workflow should be something like:
npx testops analyze --input test-results.json
Enter fullscreen mode Exit fullscreen mode
Then:
npx testops report
Enter fullscreen mode Exit fullscreen mode
And you get a dashboard.
The dashboard focuses on the questions developers actually care about:
How many tests do we have?
How many are failing?
Which ones are flaky?
Which ones are slow?
What happened in previous runs?
Enter fullscreen mode Exit fullscreen mode
Testing the tester
One of the most important parts of building a developer tool is testing it against something real.
A synthetic demo can prove that the software works in theory.
It doesn’t prove that developers can use it in their projects.
So the validation workflow is:
Existing project
↓
Existing tests
↓
Generate real test report
↓
Feed report into TestOps
↓
Compare metrics
↓
Create intentional flaky test
↓
Run repeatedly
↓
Verify detection
↓
Create slow test
↓
Verify performance detection
↓
Modify snapshot
↓
Verify regression detection
↓
Run inside CI
Enter fullscreen mode Exit fullscreen mode
This is the test that matters.
Not whether the landing page looks good.
Not whether the dashboard has impressive charts.
Whether it can survive a real repository.
What I learned
The interesting part of test infrastructure isn’t necessarily generating more tests.
It’s understanding the tests you already have.
A team can have 5,000 tests and still have a reliability problem.
More tests don’t automatically mean more confidence.
Sometimes the real question is:
Which tests can we trust?
That’s the problem TestOps Kit is designed around.
The goal
The long-term idea is bigger than a CLI.
A mature version could understand:
- Test history
- CI environments
- Failure patterns
- Code changes
- Test ownership
- Dependency relationships
- Visual changes
- Performance regressions
- Failure clusters
Eventually, the system could answer questions like:
Which tests became unreliable after this deployment?
Or:
Which changed files are responsible for most of the tests we need to run?
Or:
Which tests have consumed the most CI time this month?
But the first version starts with something much simpler:
Measure test reliability instead of guessing about it.
Because once developers stop trusting their test suite, the test suite has already become technical debt.
And that’s a problem worth measuring.
TestOps Kit — Flaky Test Detection, Test Reliability & CI Toolkit
TestOps KitProduction Test Reliability Toolkit for Modern JavaScript TeamsYour test suite is supposed to give you confidence.Instead, you get: Random CI failures Tests that pass locally but fail in CI Slow test suites Repeated failures nobody investigates Visual regressions Hundreds of test results to manually inspect Flaky tests that waste hours every week TestOps Kit turns raw test results into actionable test reliability data.It analyzes your existing test reports, tracks test history, detects flaky tests, identifies slow tests, generates quarantine manifests, checks visual snapshots, and produces a standalone reliability dashboard.No need to replace your existing test framework.What you getFlaky Test DetectionTrack test behavior across multiple runs and identify tests that repeatedly switch between passing and failing.See: Flakiness percentage Failure frequency Run history Test duration Failure information Test Reliability DashboardGet a single view of your test suite: Total tests Pass rate Failed tests Flaky tests Slow tests Test duration Historical runs Failure details Test QuarantineGenerate a quarantine manifest for tests that exceed your configured flakiness threshold.Instead of manually maintaining a list of unreliable tests, let TestOps identify them from actual test history.Slow Test DetectionFind tests that are consuming disproportionate CI time.Use duration history to identify tests that deserve optimization.Visual Regression ChecksCreate deterministic snapshot baselines and detect unexpected changes using content hashing.Useful for: HTML snapshots JSON snapshots Generated assets Test fixtures Visual regression workflows CI ReadyDesigned to fit into existing CI pipelines.Use it with GitHub Actions and your existing test infrastructure.Multiple Report FormatsThe toolkit supports common machine-readable test reports including: Playwright JSON Vitest/Jest-style JSON JUnit XML Test Impact AnalysisIdentify candidate tests related to changed files so you can focus testing where changes actually occurred.Stress TestingRun the same command repeatedly to expose intermittent failures.Example:npx testops stress –command “npm test” –runs 20 Developer CLIEverything is accessible from the command line.npx testops analyze –input test-results.json npx testops report npx testops quarantine npx testops snapshot snapshots –strict npx testops stress –command “npm test” –runs 20 Built for JavaScript developers TypeScript developers SaaS teams Freelancers Agencies QA engineers DevOps engineers CI/CD-heavy projects Teams using Playwright, Vitest or Jest Requirements Node.js 18+ npm Git Existing automated test reports Included TestOps CLI Reliability analyzer Flaky test detector Test history Quarantine generator Slow-test analysis Stress runner Visual snapshot checker Git impact analysis Standalone dashboard GitHub Actions integration Demo fixtures Complete documentation Example configuration MIT license Why TestOps?Test generation is only half the testing problem.The other half is knowing whether the tests you already have can actually be trusted.TestOps focuses on that problem.Know which tests are stable. Know which ones are flaky. Know which ones are slowing CI down.Build with confidence instead of debugging CI noise.
boukataya.gumroad.com