ooligo

Codility

technical-assessment coding-assessment · technical-screening · live-coding-interviews · skills-intelligence
MCP API
Recruiting & TA
7.6 /10

What it is

Codility is a technical-assessment platform for engineering hiring, and the third pole next to HackerRank and CodeSignal — take-home coding screens, live interviews inside a real VS Code environment, and a workforce-skills module that points the same tasks at people you already employ. It has been doing this since 2009 and raised a $22M Series A in January 2020 led by Oxx and Kennet Partners, its first institutional round after a decade on its own revenue. No round since, and the product shows it: depth on assessment methodology, no sprawl into adjacent HR software. The customer wall names GitHub, SpaceX, Tesla, EY, LSEG, SAP, Barclays and Citi.

Three modules. Screen runs asynchronous assessments against a library the vendor puts at 1,300+ tasks, each written by assessment engineers, reviewed by occupational psychologists, and tested for AI solvability. Interview is live coding in VS Code with recording, transcripts and a shared scoring framework. Skills Intelligence applies the same tasks to current engineers to map capability and staffing gaps rather than to filter candidates. Codility carries SOC 2, ISO 27001, CCPA/GDPR and WCAG 2.2 AA claims on its own site.

Why it shows up in recruiting stacks

  • It is the only one of the three majors that publishes a price. HackerRank and CodeSignal both put their entire pricing behind a form — checked again today, neither page shows a figure. Codility publishes $1,200/yr and $6,000/yr and lets you sign up without talking to sales. For a team that has to start assessing next week, that is the whole argument.
  • Custom tasks get authored from your own codebase over MCP. Codility ships an MCP server — “Build with MCP” in the Tasks menu — so an MCP-compatible coding agent points at your repo, a pull request, or a job description and writes a validated multi-file task with test cases, published straight to your task library. It needs admin permission, and it starts at the Scale tier. Neither HackerRank nor CodeSignal has an MCP path. On the Custom tier Codility names the Claude Code CLI directly, with the candidate’s agent activity reviewable per assessment.
  • AI use in the assessment is a dial you set. Codility’s position is that the question is not whether a candidate uses AI but whether they can build with it, so it sells AI-specific tasks and per-assessment visibility into what the candidate asked the assistant. The agentic AI copilot appears from Starter; Cody, the candidate-facing chatbot, from Scale.
  • Integrity targets identity, not tooling. Identity verification, impersonation detection, plagiarism detection and proctoring signals — with a desktop app on the Custom tier that flags unauthorized applications, including ones marketed as undetectable. That is the right threat model for 2026: see AI interview fraud for why the proxy interviewer, not the copilot, is the expensive failure.
  • Scoring is rules-based, and that matters legally. Codility describes auditable scoring through its evaluation engine plus automated code analysis for quality, maintainability and complexity, with the AI features sitting on the candidate’s side of the assessment rather than in the decision. If your counsel is working through Illinois HB 3773 or NYC LL 144, that distinction is the first thing to confirm — in writing, from the vendor, for the exact configuration you buy.

Pricing reality

  • Starter — $1,200/yr, annual billing only. 120 invite credits a year, 1 platform user, unlimited collaborator seats, a limited task library, plagiarism detection, basic proctoring with candidate snapshots, self-serve support. Works out to $10 per invite.
  • Scale — $600/mo, or $6,000/yr billed annually (two months free). 300 credits a year capped at 25 per month, 3 platform users, a broader library spanning engineering, AI and system design, and MCP task authoring. $20 per invite annually, $24 monthly.
  • Custom — quote only. Nearly everything that makes an assessment programme defensible lives here: ATS integrations (Greenhouse, Lever, Ashby, Workday, SAP), SSO and SAML, full API, weighted scoring, premium video proctoring with ID verification, the desktop app, interview recording and transcripts, Code Health, Skills Intelligence, the business task libraries, and validity plus adverse-impact studies from in-house occupational psychologists.

One invite credit covers one candidate in either Screen or Interview, so the meter is candidates, not seats. Scale’s 300 credits support roughly 8–12 engineering hires a year at an approximate 25:1 screen-to-hire funnel. What pushes teams to a quote is usually the 3-seat and 25-per-month caps rather than any missing capability.

Best for

Engineering and TA leads hiring 5–30 engineers a year who want structured, defensible technical screening on a card this week instead of a procurement cycle — and who are willing to review submissions themselves, because nothing here auto-ranks candidates for them.

Not for you if scheduling is the bottleneck rather than screening quality — buy GoodTime or Modernloop instead. Not for you if ATS write-back is non-negotiable on a self-serve budget: at Starter or Scale, results live in Codility and someone copies the outcome into your ATS by hand. And not for non-engineering hiring, where TestGorilla covers far more role types for $75/mo.

Versus the alternatives

HackerRank is the volume leader and the safer answer above roughly 50 engineering hires a year — bigger library, more recruiter familiarity, quote-only pricing that lands in the low-to-mid five figures annually. Pick it when the assessment programme is already staffed and the constraint is throughput. CodeSignal is the pick when you want a certified evaluation score that transfers across roles and req cycles, and its detection stack is the closer comparison on cheating; HackerRank vs CodeSignal settles that pair. micro1 is the fastest-growing entrant and a different purchase: Zara conducts the interview instead of equipping your interviewer. The rule is clean — Codility measures candidates for humans to judge, micro1 replaces the human in the room. Pick micro1 when interviewer hours are the constraint; pick Codility when question quality and auditability are.

Watch-outs

  • Almost every compliance-grade feature is Custom-gated — SSO, API, ATS write-back, ID verification, the desktop app, fairness studies. Guard: list which of those your security review and ATS actually require before you buy self-serve. If even one is on the list, go straight to the quote; retrofitting the integration later means re-running the rollout.
  • Credits are annual and Scale caps at 25 a month. Guard: size from last year’s peak screening month, not the average. A January req spike burns a 300-credit allowance well before Q3, and the cap blocks the obvious workaround.
  • Task leakage is a live risk on any published library. Codility removes leaked tasks and says so on every tier. Guard: rotate tasks per req cycle and author 2–3 MCP-built tasks from your own repo for final rounds, where leakage costs the most.
  • No AI scoring means humans still read every submission. Guard: budget review hours before rollout — roughly 10–15 minutes per submission — and name the reviewer per req, or the queue becomes the new bottleneck and screens go unread.
  • Skills Intelligence, the reason to prefer Codility over a pure screener, is Custom only. Guard: if internal capability mapping is the business case, do not pilot on Starter; the pilot cannot demonstrate the thing you are buying.