ooligo
claude-skill

Audit an AI interviewing stack against US employment-AI rules

Difficulty
advanced
Setup time
90min
For
talent-acquisition · recruiter · legal-ops-manager
Recruiting & TA

Stack

A Claude Skill that takes your requisition list, your tool ledger, and your evidence pack, and returns a control matrix showing which employment-AI obligations attach to which req, which controls you can actually prove, and which gaps accrue penalties per day. It grades on artifacts you supply — a published bias-audit URL, a consent timestamp export, a deletion runbook — not on vendor claims, and it never emits a compliant/non-compliant verdict. The output is the thing you hand counsel so their first hour starts at the judgment calls instead of at “which tools do you use, and where?”

When to use

  • You are turning on an AI interviewer, an AI video screen, or resume ranking, and someone asked what the exposure is before it goes live.
  • A req just opened in a jurisdiction your stack has not served before. Remote-eligible reqs are the usual trigger: a single “US-wide” posting can attach New York City, Illinois, and California obligations at once.
  • The annual NYC bias-audit renewal is approaching and you need to know which tools are in scope before booking the auditor — auditor engagements are quoted per tool.
  • A candidate asked how the decision was made, or requested deletion of their interview, and nobody is sure what the runbook is.

When NOT to use

  • To produce the bias audit itself. NYC Local Law 144 requires an independent auditor with no employment or financial relationship with the employer that would compromise independence. A skill you run cannot be that auditor. This skill checks whether an independent audit exists, is within its clock, and is published in the required form.
  • As a sign-off. There is no aggregate score in the output, deliberately. A number invites someone to screenshot it into a board deck as a clean bill of health.
  • After a charge or demand letter lands. At that point the artifact set is discovery material and counsel drives. Generating a parallel internal assessment of the same facts creates a document you did not need.
  • On stacks with no automated scoring. If a human reads every application and nothing ranks or filters, most of the matrix does not attach.
  • Outside the US. The bundled matrix covers federal, state, and city rules only. EU AI Act Annex III employment duties are a different scope and would need their own reference file.

Setup

  1. Drop the bundle at apps/web/public/artifacts/ai-interview-compliance-audit-skill/SKILL.md into your Claude Code skills directory, with the references/ folder alongside it.
  2. Have counsel review the matrix once. references/1-jurisdiction-matrix.md holds every statutory parameter the skill grades against — the 10-business-day NYC notice window, the 30-day Illinois destruction window, the four-year California retention floor. One review pass, then keep the checked: date current. The skill refuses to grade if that date is more than 90 days old.
  3. Fill the ledger. references/2-tool-and-stage-ledger.md wants reqs with work locations, and one row per tool that touches a candidate between application and offer. Include the tools you did not buy for screening — ATS match scores and sourcing fit scores are the ones teams forget, and they rank.
  4. Record review capacity. Part C of the ledger asks how many applications a human actually reviews per req. “400 applicants, recruiter works the top 40” is the fact that turns a ranking tool into a functional filter, whatever the configuration screen calls it.
  5. Assemble the evidence pack per references/3-evidence-pack-index.md. Fifteen artifact rows; write MISSING where it is missing. That first pass is usually where the real finding shows up.

What the skill actually does

Six steps. Scope resolves before any control is graded, because the same configuration is lawful in one state and a per-day violation in another — and because the most common real-world error is scoping the audit to where the company is headquartered rather than to where candidates apply from.

  1. Resolve scope. Build a req × jurisdiction × stage matrix. If a req has no work-location list, the skill stops and asks rather than inferring from the company address. Reqs open to states the evidence pack has no artifacts for come back as gap, not not-applicable.
  2. Classify each tool from configuration. Four questions: does the output gate a stage transition at any threshold; does it sort a list a recruiter works top-down; does a human see the score before making the call; does it analyze video or audio for characteristics used to evaluate fitness. The vendor’s own label goes in a separate column and is never the answer. The Local Law 144 test turns on whether the output substantially assists or replaces discretionary decision-making, which is a fact about your setup, not about their product — and the vendor has an incentive to read it narrowly.
  3. Grade each control on evidence, with a citation. Four statuses: evidenced (artifact cited verbatim, with the passage that does the work), unevidenced (probably fine in practice, nothing to show), gap, and counsel-review. There is no compliant. A control with no citation cannot be evidenced — that rule is what stops a fluent report about documents nobody has.
  4. Run deterministic date checks. Bias-audit clock, flagged at 10 months rather than 12 so there is runway to book the auditor. Notice lead time counted in business days against the tool’s first run. Illinois consent captured before the interview, not bundled into a post-interview acknowledgment — ordering is the entire control. Deletion runbook turnaround against the 30-day window, and whether it names downstream recipients and backups or only the primary ATS.
  5. Order remediation by exposure. New York City counts each day an in-scope tool runs out of compliance as a separate violation, and each missed candidate notice as its own separate violation, at up to $500 for a first violation and $500 to $1,500 for subsequent ones. A notice gap on a 400-applicant-per-month req therefore compounds against a fixed remediation cost; the same gap on a paused req does not. The list sorts on that difference.
  6. Emit the report. Matrix, deterministic results, remediation, then the counsel-review items verbatim with both readings stated.

What changed recently, and why the matrix is a file

Colorado is the reason the statutory parameters live in a dated reference file instead of in the model’s head. SB 24-205 — the 2024 Colorado AI Act, with its risk-management programs and annual impact assessments — never took effect. Its start date moved from February 2026 to June 2026, and then SB 26-189, signed 14 May 2026, replaced it with a narrower notice-and-transparency framework effective 1 January 2027: pre-use notice, a plain-language disclosure within 30 days of an adverse outcome, a right to request human review, and a three-year records floor. Any checklist still grading impact assessments is auditing a repealed statute. The Colorado AI Act explainer walks the replacement in full.

Illinois moved too. The AI Video Interview Act has run since 2020, but the Human Rights Act amendment banning discriminatory AI effects — including zip code as a proxy for a protected class — took effect 1 January 2026, and the Department of Human Rights notice rules are still in rulemaking after a withdrawn first draft. California’s FEHA automated-decision-system regulations have applied since 1 October 2025, with a four-year records floor that points the opposite way from a candidate’s deletion request. The skill surfaces that tension rather than resolving it.

The December 2025 executive order directing a DOJ task force to challenge state AI laws is in the matrix as context, not as a control. No state rule above has been displaced. Treat it as a reason to keep controls documented and portable, not as a reason to retire a row.

Cost reality

  • Per audit run — roughly 40-80k input tokens (matrix, ledger, notice texts, vendor docs) and 6-10k output. At Claude Sonnet list rates that is about $0.30-0.60 per run. Estimate, from the token shape of a three-req, five-tool stack.
  • Setup — 90 minutes, and that number is honest only if the evidence pack already exists somewhere. Teams assembling it for the first time spend 4-8 hours, most of it finding out that a control everyone assumed was handled has no artifact behind it.
  • Counsel time saved — an outside employment-counsel first pass on a multi-jurisdiction stack runs 8-20 hours at $350-700/hour, much of it intake: what tools, which reqs, which states, what does the score do. Arriving with a filled ledger and a graded matrix moves that intake in-house. The judgment calls still bill.
  • What it does not save — the independent bias audit. That is a separate engagement, quoted per tool, and this skill cannot substitute for it.

Success metric

  • Unevidenced count trending to zero. The first run’s split between evidenced and unevidenced is the real baseline. Controls moving from unevidenced to evidenced without any operational change is the intended outcome — it means the artifact now exists.
  • Bias-audit clock never past 10 months. A deterministic check that should never fire twice on the same tool.
  • Notice lead-time distribution, not average. Track the fifth percentile of business days between notice and first tool run per req. Averages hide the fast-moving reqs, which are exactly where the lead time fails.
  • Time from new-jurisdiction req to graded matrix. Should land under a day once the ledger exists.

vs alternatives

  • vs a spreadsheet checklist. The status quo, and it fails on two specific things: it scopes by company location rather than per req, and it has no clock, so a bias audit silently ages past 12 months while the tool keeps running. The skill’s step 1 and step 4 exist because of those two failures.
  • vs vendor compliance attestations. HireVue, Sapia.ai, and others publish bias-audit summaries for their models. Useful, and they satisfy vendor-side rows. They do not satisfy yours: the Local Law 144 duty sits with the employer or employment agency, and the audit that matters covers your configuration and your candidate pool. The evidence index keeps vendor and employer artifacts in separate columns so this cannot blur.
  • vs a dedicated audit vendor. Firms that perform independent AEDT bias audits do the thing the skill structurally cannot. Use the skill as the readiness pass before you engage one — arriving with a classified tool ledger cuts scoping back-and-forth — and as the between-audits recheck when a req opens in a new state.
  • vs asking counsel to run the whole thing. Correct for the judgment calls, expensive for the inventory. The split this workflow proposes: you own the ledger and the evidence pack, counsel owns the matrix review and the counsel-review queue.

Watch-outs

  • The model states a legal conclusion. Guard: the status vocabulary has no compliant/non-compliant value and there is no aggregate score. Judgment calls route to counsel-review with both readings printed.
  • Statutory parameters drift. Guard: thresholds live in references/1-jurisdiction-matrix.md with a checked: date, and the skill refuses to grade past 90 days. This landscape moved three times between August 2025 and May 2026.
  • A vendor attestation gets counted as your compliance. Guard: row 14 of the evidence index is marked vendor-side and cannot satisfy the employer’s audit or publication rows.
  • Scoping to headquarters. Guard: step 1 will not proceed without per-req work locations, and remote-eligible reqs expand to every state in the accepted list.
  • Classification laundering. Guard: classification comes from configuration facts in the ledger, with review-capacity data attached. A tool that ranks 400 applicants where a human reviews 40 is doing decision work whatever its label says.
  • unevidenced read as a pass. Guard: those rows sort into remediation alongside gaps, with the missing artifact named.
  • Retention and deletion pulling opposite ways. Guard: the California four-year floor and an Illinois deletion request collide on the same record. Step 4 flags the conflict; counsel decides it once and the report cites that memo.

Stack

The bundle lives at apps/web/public/artifacts/ai-interview-compliance-audit-skill/ and contains:

  • SKILL.md — the skill definition
  • references/1-jurisdiction-matrix.md — statutory parameters with a checked: date
  • references/2-tool-and-stage-ledger.md — fillable req, tool, and review-capacity inventory
  • references/3-evidence-pack-index.md — control-to-artifact mapping plus notice and consent scaffolding

Assumes Claude for the run. The stack under audit typically includes an ATS such as Greenhouse and one or more AI screening tools — HireVue, Sapia.ai, or similar.

Related reading: NYC Local Law 144, AI screening bias, AI resume screening, AI policy for recruiting teams, candidate experience.

Files in this artifact

Download all (.zip)