Product overview

How Litmus works.

Evals for humans.

What Litmus does

Litmus generates and hosts coding interviews built from your codebase, tickets, and job descriptions. Instead of generic OAs and take-homes, candidates build a real feature with their usual tools while we track file iteration, commit history, and AI usage. A round can run async, or live on a call with your team. Every submission runs against a grading harness in a sandbox, and you get an evidence-based report on how each engineer builds.

Where Litmus fits in

Litmus can fit in anywhere in your hiring pipeline. Here are a few places other companies have found Litmus useful:

ApplicationRecruiter screenOnsiteOfferTop of funnelFirst roundOn-site
  • Top of funnel · Submitted with the application: a technical read before a resume even reaches a human.
  • First round · The most common placement: narrow a large inbound pool before anyone spends onsite or oncall time.
  • On-site · For work trials and extended projects: granular tracking across a long, complex task.

What changes for your team

Engineer time

Zero

engineering hours to run your top of funnel

No designing interview problems, no screening calls, no grading early submissions. By the time an engineer meets a candidate, they've already built in your context.

Your workflows

Plug-in

to the process you already run

Candidates pass to and from your ATS automatically, or you can manage them start to finish on Litmus. Nothing about your pipeline has to change.

Signal

Work-trial

depth from every submission

How someone builds across a long task, how they use AI in practice, how they explain their decisions. None of it shows up in a standard interview.

How assessments are built

Host an existing assessment you love, or generate one from scratch by linking your repos, tickets, and job descriptions. Test candidates on real end-to-end feature work to see how they build. Engineers solve problems like they would on the job, not just implement solution specs.

Build me an assessment from this ticket

ENG-482Bulk export pipeline

Litmus building now · ~2 minutes

ReadENG-482 + your stack
WriteREADME.md
Scaffoldexports/ · 14 files
Harness6 graded requirements
Questionswalkthrough · 4
Exports150 min · TypeScriptReview

Real assessments built this way.

MarketGraded belowA prediction-market engine with an automated market maker, graded under concurrent trades.
CalA booking engine that survives DST transitions and concurrent hold stampedes.
BlocksA paged KV-cache scheduler for LLM serving: prefix sharing, copy-on-write, preemption.
FlashA streaming attention kernel: online softmax, no N×N matrix, stable at large magnitudes.

What the candidate experience looks like

01

Initialize

Setup is two terminal commands. The assessment pulls straight into the candidate’s own workspace.

How live interviews work

Illustration of a live round in progress: the candidate's shared screen with their editor and test output, and the Litmus console beside it showing the session in progress, the notetaker in the call, and grading waiting for the transcript.

Litmus also supports live technical interviews: same problem, same repo, same submission, same analysis. The candidate books a time with your team and works in their own setup while the interviewer watches.

In a normal live interview, the interviewer juggles conducting the assessment, taking notes, and forming a verdict, and their notes are the only record. With Litmus, every command, tool, and prompt is captured automatically, the submission runs the same checks as any take-home, and the transcript lands in the report.

The interviewer gets to spend the hour on the candidate.

How grading works

Grading is tailored to each assessment and designed to surface signal on how an engineer would actually build on your team. The harness replays each submission against scenarios it has never seen, the grader reads the whole working record, and the result is one defended judgement on a 5-point scale, with every claim cited to evidence.

Run by the harness

Hard requirements replay in a live sandbox

OutcomesConcurrencyInvariants

Read by graders

The whole working record, judged with evidence

CodeAI conversationWalkthrough
Both land in one evidence-taggedE7E7 report
Interactive example of a Litmus candidate report: an overview with score and engineering profile, the candidate's AI prompt log, their submitted code, and a playable walkthrough recording, in the same tabbed layout reviewers use.

Caroline Smyth

4.2/ 5.0Rare
9
Outcomes passed
of 11
2
Outcomes failed
of 11
4
AI prompts
2
Commits
this session
98
Files touched
30m
Session
Dimension scores
Correctness
5
Judgement
4
Decomposition
5
Ownership
4
Verification
5
Understanding
4
Summary

The whole session ran about thirty minutes.

Caroline sent one prompt asking the assistant to read the brief and frame an architecture, then directed a multi-agent build against a shared API contract, sent verifier agents across the functional and atomicity domains before accepting the plan, probed the overdraw path directly, and closed by committing the result. Direction, verification, and the hard questions came from the candidate; the agents did the typing.

Profile
Caroline ran this as an orchestration exercise: two short prompts set up a contract-first, multi-agent build with verifier agents, and the harness bears the result out. The guarded-debit IMMEDIATE transaction on the trade path held through three concurrent bursts, and the double-resolve guard kept payouts exact.The replay passed nine of eleven checks outright: market creation, AMM-exact quoting, end-to-end trading with conserved balances, overdraw rejected with state intact, and correct resolution. The two that failed are both pricing edges: bounded micro-trade rounding drift and a raced quote that fills at the shown price rather than requoting, a limitation Caroline called out unprompted in the walkthrough.The walkthrough walks the trade path from route to guarded debit to position upsert and explains why resolution is a guarded status flip, matching the code as submitted.

Integrations

GitHubLinearJiraNotionGreenhouseAshbyLever

Questions? Email us at founders@litmushiring.com