Threat index · Q3 2026 · Updated April 2026
AI Cheating Threat Index: Q3 2026
Quarterly tracking of commercial AI overlays, open-source forks, on-device LLMs, remote-access tools, and proxy exam services, with threat scores, detection verdicts, and defense effectiveness ratings. Updated at the start of each quarter.
Q3 2026 · Data through 2 September 2026 · Next revision: December 2026
Executive summary
Q3 2026 in three sentences.
Every price in this edition was read from the vendor's own page in the first week of September 2026, and several of our Q2 figures did not survive that check. The composite Threat Score, a weighted average across categories, reached 89/100 in Q3 2026, up from 87/100 in Q2. Corrections to our own prior edition are listed in the delta section.
The adversary profile remains bifurcated: the casual cheater buys Cluely; the determined cheater compiles their own binary. What changed in Q3 is the second path. Five of the ten actively-maintained open-source projects now document a fully offline mode, and one of them shipped that capability during this quarter. A candidate on that path generates no outbound traffic to any AI endpoint, because there is no endpoint.
Defense effectiveness did not move at all. We read the quarter's release notes for Zoom, Google Workspace, Webex, macOS and Windows: no vendor shipped anything that detects or blocks a capture-excluded overlay. Six of seven monitored detection approaches remain bypassed, and network-layer enforcement remains the only approach in this index with no documented bypass.
Jump to section
- Executive summary
- Composite threat score
- 1. Commercial AI overlays
- 2. Open-source forks
- 3. On-device LLMs
- 4. Remote-access tools
- 5. Proxy & exam fraud services
- 6. Hardware attack surface
- 7. Defense effectiveness ratings
- 8. Quarter-over-quarter delta (Q2 → Q3)
- 9. Scoring our Q3 forecasts
- Methodology & how to cite
Composite threat score
89/100. Critical. Up 2 points from Q2.
Weighted average across seven threat categories. Higher score = harder to defend against with current detection-only architectures. A score of 100 would represent a landscape where no deployed defense has any effectiveness.
Two scales, not one: categories below (and the 87 composite) are scored 1–100. Individual tools in the sections that follow are scored 1–10. Full weighting in Methodology.
Commercial AI overlays
Open-source projects (compiled)
On-device LLMs
Remote-access tools (proxy use)
Proxy & exam fraud services
Hardware (earpieces, smart glasses)
Browser automation / scripting
Composite: all categories
Scores reflect detection difficulty for current deployed assessment defenses, weighted by observed prevalence and ease of use for the median attacker. Scores are not a measure of absolute harm; a 100 would represent a theoretically undefeatable threat ecosystem.
§1
Six commercial tools. Prices up sharply, invisibility now sold separately.
Commercial invisible overlay tools are subscription software products with pricing tiers, customer support, and changelog updates. Their viability depends on continued access to LLM APIs, which is also their single point of failure under network-layer enforcement.
Every price below was read from the vendor's own pricing page on 2–3 September 2026. Four of the six had moved since our Q2 edition, in one case by a factor of eight. Where a vendor states two different prices on its own site, both are shown.
Reading the scores: individual tools below are scored 1–10 (threat score). Categories and the overall index are scored 1–100. The two scales aren't on the same axis; see Methodology for the weighting formula.
| Tool | Price | Threat score | Detection difficulty | Key capability | Network dependency | Q3 status |
|---|---|---|---|---|---|---|
| Cluely | Free · $19.99/mo Pro · $149.99/mo Pro + Undetectability | 8/10 | HARD | Undetectability moved into a separate top tier; 12+ languages, stated 300ms response | Required: cloud LLM | ACTIVE · repriced |
| Interview Coder | $799 one-time · $299/mo | 8/10 | HARD | Windows support added; 150,000+ claimed users, up from 100k+ | Required: LLM backend call per query | ACTIVE · +8× price |
| Parakeet AI | $149.90/mo · $78/week · $599.90/yr | 9/10 | HARD | Real-time audio + LLM, 50+ languages; claims invisibility in dock and Activity Monitor | Required: transcription + LLM calls | ACTIVE |
| Ultracode AI | $799 (homepage and checkout) · $899 (four subpages) | 8/10 | HARD | Dropped its Q2 taskbar-visibility caveat; now claims zero trace on Windows and macOS | Required: cloud LLM backend | ACTIVE · claim hardened |
| LockedIn AI | $109.98/mo list ($54.99 promo) · $1,999 lifetime ($999.50 promo) | 9/10 | HARD | Duo adds live remote control of the candidate machine by a third party | Required: AssemblyAI, Azure OpenAI, ElevenLabs, Pinecone | ACTIVE · new capability |
| Final Round AI | $150.00/mo · $25/mo billed annually | 7/10 | MEDIUM | Stealth scoped to screen share and recording only; no taskbar or process claim made | Required: transcription and LLM | ACTIVE |
Every LockedIn AI price is displayed under a countdown timer; the list prices are shown first and the promotional prices in brackets. Ultracode AI states $799 on its homepage and in its checkout URL, and $899 on four product subpages; we report both rather than choose.
Four new entrants, and all of them arrived through the App Store
The six tools above are the ones we track in depth. Four more launched inside this quarter, and how we found them is itself the finding: none appeared on Product Hunt or Hacker News. We reconstructed 176 Product Hunt launches across July and August, and separately censused all 8,220 Show HN posts in the window. Between them they surfaced exactly one covert interview assistant — a browser-based tool that explicitly disclaimed stealth (“not a stealth tool”) and whose domain is now parked for sale. All four genuine entrants were found instead in Apple's App Store metadata, which publishes a real first-release date.
The same invisibility holds in the press. Across the quarter we found no named-tool story about interview cheating in any tier-one technology outlet — TechCrunch, The Verge, Ars Technica, Wired, CNBC and 404 Media were each swept and returned nothing in the window. The products that were named were named in HR trade press and in contributor columns, several of which sourced their prevalence figures from vendors that sell detection. A buyer relying on mainstream coverage to know what exists in this category will not learn it there.
| Tool | First released | Price | Platform | Invisibility claim |
|---|---|---|---|---|
| YIVIDIA Live AI Assistant | 2 July 2026 | $3.99/month, or credit packs from $2.99 | macOS, iOS, iPadOS | None. Always-on-top floating panel |
| ghostmind | 6 July 2026 | Free; requires the candidate's own OpenAI key | macOS 14+ | None claimed; ships full coding-interview cards |
| Sotto Interview Copilot | 13 July 2026 | $7.99/month · $49.99/year | iOS, iPadOS | None. Deliberately a second-device design |
| IntervuMate Interview Copilot | 27 August 2026 | $9.99/week · $24.99/month · $94.99/year | iOS, macOS | “Completely hidden from all Screen Share Applications like Teams, Zoom, Meet” |
We are not scoring these four individually yet. Three make no invisibility claim at all, none publishes a verifiable user count, and one is a version 1.0 released a week before this edition closed. They move the commercial category by a single point, on ease of access rather than on capability, and we will revisit them in Q4 with usage evidence rather than launch dates.
One near-miss we could not date. SkillCopilotAI advertises “completely invisible and undetectable in any online meeting software and online coding test platforms” at $28.90/month billed yearly, and names HireVue among its targets. Its first public trace is 14 August 2026, but its domain was registered in March, so we cannot establish it as an in-window launch and have left it out of the count.
§2
Ten active projects, not twenty. And none of them is a fork.
The open-source ecosystem is the existential challenge for detection-first proctoring. Each project can be compiled with a completely custom binary name, icon, and process signature, rendering signature-based detection permanently obsolete for any candidate with basic developer skills, which describes most of the candidates being assessed.
The smaller number is not good news. Between them these ten projects have been downloaded 410,506 times as compiled binaries, and the two largest projects by stars are both dormant — meaning the active ecosystem sits behind an even larger body of abandoned but still-working code. Counting by GitHub topic tag inflates the total: one repository carrying the category topic and 252 stars contains a single 1,275-byte README and no code at all.
| Fork | Platform | Threat score | Distinguishing capability | Process-name evasion | Network dependency |
|---|---|---|---|---|---|
| OpenCluely | GitHub (MIT) | 9/10 | Invisible overlay for DSA/coding; multi-language; Gemini integration | Full: compile with any binary name | Required: Gemini/OpenAI API |
| Pluely | GitHub (Tauri/Rust) | 9/10 | 10 MB; 50% less RAM than Cluely; invisible in Zoom/Teams/Meet; GPT-4/Claude/Gemini/Grok multi-model | Full: Rust source, recompile in <1 hr | Required: multi-model API support |
| Natively | GitHub (MIT) | 10/10 | Local RAG; BYOK; zero server storage; explicitly disguises as Terminal, Activity Monitor, or System Settings | Full: documented feature, named as system utilities by design | Optional: BYOK supports local inference |
| MindWhisperAI | GitHub (MIT) | 9/10 | GPT-4o/Claude/Gemini/Grok support; stealth mode; handles coding, system design, behavioral | Full: MIT license, no telemetry, fully forkable | Required: multi-API |
| ShadeCoder | GitHub | 8/10 | Whisper STT integration; screen-capture → code pipeline; low latency vs. Cluely | Full: open source | Required: transcription + LLM |
| LeetcodeWizard | GitHub | 8/10 | LeetCode-specific; includes humanizer pipeline targeting perplexity normalization to defeat AI detectors | Full: open source | Required: LLM + humanizer API |
| DIY (Tesseract + Whisper + any LLM) | Any developer | 10/10 | No GitHub signature exists; OCR + STT + LLM in a few hundred lines of Python; fully custom | N/A: no signature | Required: LLM API |
| DIY with local model | Any developer | 10/10 | No external network trace; Ollama-backed; zero internet required at exam time | N/A: no signature, no network | None; fully offline |
§3
A 30B multimodal model, Apache-2.0, in a 16.8 GB file.
Local LLM inference remains the threat vector that network enforcement alone does not reach. A candidate running Ollama, LM Studio, or a self-compiled inference server generates zero external network traffic. Their device calls no AI API. DNS queries to known AI providers are irrelevant, because none are made.
Through Q2 the practical constraint was hardware. Q3 removed a large part of it. On 10 August 2026 Meta released Muse Glimmer 30B under Apache-2.0 — a 30-billion-parameter dense, natively multimodal model with a 128K context window. Its 4-bit quantisation is a 16.76 GB file, and Meta's own description is that it is “small enough to run on a Mac or PC with a single consumer GPU.” It has been downloaded more than 620,000 times.
At the other end of the scale, Liquid AI's LFM2.5-2.6B ships a 1.59 GB quantisation the vendor reports running at 220 tokens/second on an Apple M5 Max and 30 tokens/second on a phone, in under 2.5 GB of memory. Note that despite being freely downloadable it is not Apache-licensed: the LFM Open License conditions commercial use on the licensee earning under $10 million a year.
| Runtime / model | Threat score | Minimum hardware | Internet required? | Q2 prevalence | Detection surface |
|---|---|---|---|---|---|
| Ollama (any 7B model) | 8/10 | 8 GB RAM + modern CPU (M1/M2 Mac, Ryzen 7) | No; fully local after download | HIGH; mainstream on developer machines | Device activity + hardware resource signals |
| LM Studio (any model) | 8/10 | 8 GB RAM; GUI installer for non-developers | No; fully local after download | HIGH; lowers technical barrier further | Device activity + hardware resource signals |
| llama.cpp (CLI) | 9/10 | 4 GB RAM for quantized models | No | MEDIUM; developer-only | Device activity |
| GPT4All | 7/10 | 8 GB RAM; very low-skill GUI | No | MEDIUM; consumer-friendly packaging | Device activity + hardware resource signals |
| Offline Gemma 2B (phone) | 9/10 | Modern Android or iOS device | No | EMERGING; ML Kit on-device API | Second device; outside candidate machine |
§4
The overlay vendors now sell remote control as a feature.
Remote-access tools used for exam fraud operate in two modes: human proxy, where a skilled operator sits the exam from another location, and AI pipeline, where a local script feeds questions to a remote LLM and injects answers. Both require the enrolled device to maintain a remote-control connection, which creates a network-observable signal.
| Tool | Threat score | Fraud use case | Network dependency | Detection surface | Q3 status |
|---|---|---|---|---|---|
| AnyDesk | 9/10 | Human proxy sits the exam; enrolled device shows blank screen or fake video feed | Required: relay server connection | Relay IP blocked at network layer; behavioral anomalies from remote operator | ACTIVE; widely used in proxy exam rings |
| TeamViewer | 8/10 | Same as AnyDesk; older, more detectable signatures | Required: relay server | Known relay IP ranges; process detection | ACTIVE; declining vs. AnyDesk |
| Chrome Remote Desktop | 7/10 | Requires Google account; less operational security for rings | Required: Google relay | DNS query to Google relay domains (blocked under exam policy) | MONITORING |
| Custom SSH tunnel + VNC | 10/10 | No known commercial signature; operator uses SSH for control, VNC for screen | Required, but tunneled through SSH to a controlled host | Behavioral; operator typically less fluent than genuine candidate | EMERGING; seen in APAC-targeted rings |
| AI pipeline over localhost | 10/10 | Local script: OCR screen → HTTP to local LLM → inject answer | None if local model; minimal if cloud | Device activity (localhost HTTP) + hardware signals | EMERGING |
§5
$200–$500 per exam, pay after passing: a mature fraud-as-a-service ecosystem.
Proxy exam fraud services operate as structured marketplaces: a client posts an upcoming exam, operators bid, and payment is released on successful completion. Pay-after-pass pricing removes financial risk for the buyer and creates strong performance incentives for operators. The ecosystem is most active in cybersecurity certifications, technical hiring assessments, and academic examinations.
| Service type | Threat score | Typical price | Dominant exam category | Detection challenge |
|---|---|---|---|---|
| Certification proxy rings (cybersecurity) | 9/10 | $200–$500 pay-after-pass | IT and cybersecurity certification programs | Remote desktop injection inside exam software; operator is often certified and has sat same exam before |
| Hiring assessment proxy (technical) | 9/10 | $100–$300 per session | LeetCode-style, HackerRank, CodeSignal | AI overlay or skilled human operator; cross-session intelligence needed to detect repeat operators |
| Academic exam proxy | 8/10 | $50–$200 | University finals and graduate admissions assessments | Remote desktop through screen-share software; camera feed sometimes replaced with pre-recorded footage |
| Deepfake identity fraud | 8/10 | Bundled with full fake application services | Video interview rounds, identity verification checkpoints | Live deepfake video generation; FBI-documented against US tech employers |
| Telegram / Discord fraud channels | 7/10 | Variable; answer leaks, shared accounts | All categories | Content sharing; hard to attribute; primary signal is answer similarity across candidates |
§6
Earpieces, smart glasses: the attack surface that software cannot reach.
Hardware-based cheating exists entirely outside the enrolled device and its network. No software agent, however deep, can detect a Bluetooth earpiece paired to a phone running ChatGPT in a candidate's pocket, or smart glasses with a camera and audio pipeline. This is the honest boundary of what software-layer assessment security can achieve.
| Attack vector | Threat score | How it works | Software detection? | Q3 status |
|---|---|---|---|---|
| Bluetooth earpiece + phone AI | 8/10 | Phone runs ChatGPT; candidate subvocalizes question; earpiece delivers answer | No; entirely separate hardware | ACTIVE; $30–50 earpiece, free AI |
| Smart glasses with camera | 7/10 | Camera captures screen; phone processes via LLM; earpiece delivers answer | No | EMERGING; Meta Raybans and equivalents |
| Second phone below webcam | 9/10 | Phone runs full AI chat app; candidate types question, glances briefly at response | No; camera proctoring can detect if calibrated for downward gaze | ACTIVE; the most common hardware vector |
| Hardware AI wristband / ring | 6/10 | Experimental; vibration-based Morse code delivery of answers | No | EXPERIMENTAL |
| Second laptop behind the primary | 8/10 | Positioned behind primary machine; candidate rotates to query AI, rotates back | No; camera may detect posture shift | ACTIVE |
§7
Seven defense approaches. Six bypassed. One without a known bypass.
Each defense is rated against the full threat landscape documented in Sections 1–6. “Bypassed” means a working, documented evasion technique exists and is accessible to any motivated candidate.
The macOS and Windows gap has inverted
The technique every one of these tools relies on is capture exclusion: marking a window so it is omitted from screen sharing and recording. On Windows this still works as designed. On macOS it no longer reliably does — Electron's own documentation now states that applications using ScreenCaptureKit “will capture your window despite calling win.setContentProtection(true), due to an intentional macOS change,” and two of the open-source projects independently describe their macOS invisibility as best-effort. The practical consequence is counter-intuitive and worth stating plainly: for this specific threat class, a macOS candidate is now somewhat easier to catch than a Windows one, which is the reverse of the usual assumption.
One more Q3 development belongs here. In August a small open-source utility appeared that does nothing but enumerate open windows and report which ones are excluded from screen capture, flagging them in red. Its author's stated rationale is that “some malware does it too, to stay out of recordings.” It is a hobby project of a few dozen stars, not a product. But it is the exact detection primitive this entire threat class depends on defeating, and it is now sitting in public.
Verdict legend: BYPASSED defense fails outright · PARTIALLY BYPASSED evasion exists but needs real effort · PARTIALLY EFFECTIVE catches some, misses a lot · NO KNOWN BYPASS holds under current threats. Full definitions in Methodology.
| Defense approach | Verdict | Catches | Misses | Why bypassed |
|---|---|---|---|---|
| Process-name signature scanning | BYPASSED | Unsophisticated users of commercial tools without renaming | Any open-source fork compiled with custom binary name; local models; hardware | Open-source forks can be recompiled with any process name in under 1 hour |
| Browser lockdown / secure browser | BYPASSED | Tab switching; copy-paste from other browser windows; basic tab-based cheating | Any OS-level process; overlay tools; local models; remote access; hardware | Overlay tools are native OS applications; browser restrictions have no authority below the browser |
| Keystroke dynamics analysis | BYPASSED | Automated script injection of pre-written answers (non-human timing patterns) | Manual transcription of AI-generated output; human proxy input | arXiv 2601.17280 (2026): manually transcribing AI output produces patterns statistically indistinguishable from genuine composition |
| Gaze / eye tracking | BYPASSED | Obvious downward eye movement toward a secondary device; absence from frame | Overlay positioned below webcam; audio-only pipelines (earpiece); mental recall | Overlay can be positioned so that reading gaze appears as forward-facing camera contact |
| LLM output similarity / perplexity scoring | PARTIALLY BYPASSED | Unmodified AI-generated answers pasted directly; obvious LLM boilerplate | Humanizer pipelines; rephrased AI output; answers adapted for specific context | Humanizer tools (LeetcodeWizard and equivalents) explicitly target perplexity normalization |
| Adaptive follow-up questioning | PARTIALLY EFFECTIVE | Candidates who cannot elaborate on AI-generated answers; basic AI users | Candidates who studied their answer before follow-up; audio pipelines that continue during verbal questions | Best current human-judgment method; incomplete coverage for prepared candidates |
| Network-layer enforcement (Aiseptor) | NO KNOWN BYPASS | All commercial overlays (require internet); open-source forks with network dependency; remote-access tools; encrypted resolver bypass attempts | Fully offline local LLMs after model download; hardware attack surface (second devices) | Per-session network enclave with approved-domain enforcement and OS-level signal detection. Offline local models and separate physical hardware are outside the enrolled device boundary. |
§8
Q2 to Q3 2026: what changed, and what we got wrong.
The first four rows are movements in the threat landscape. The last three are corrections to figures we published in the Q2 edition and could not reproduce when we checked them against primary sources this quarter.
| Development | Q2 status | Q3 status | Direction | Significance |
|---|---|---|---|---|
| Price of invisibility | Cluely undetectability tier at $75/month (April snapshot) | $149.99/month, base tier unchanged at $19.99 | ↑ ESCALATING | The premium for invisibility went from 3.75× to 7.5× the base plan. Vendors now price undetectability as the product rather than as a feature of it |
| Interview Coder reach and price | $100 lifetime, macOS only, 100k+ claimed users | $799 lifetime or $299/month, Windows added, 150,000+ claimed users | ↑ ESCALATING | An eightfold price increase alongside platform expansion and a claimed 50% user growth. Demand in this category is not price-sensitive |
| Remote control as a retail product | Proxy rings improvised with AnyDesk, TeamViewer, VNC | LockedIn AI sells hands-on remote control from $54.99/month | ↑ ESCALATING | The proxy threat is now a consumer subscription rather than an organised ring. Controls scoped to criminal infrastructure are aimed at the wrong adversary |
| Local inference inside the tools | Possible via Ollama, but not shipped in the overlays themselves | 5 of 10 active projects document an offline mode; one shipped it in-quarter | ↑ ESCALATING | Paired with a 30B multimodal model arriving as a 16.76 GB Apache-2.0 download, this is the path on which egress logging sees nothing |
| Platform and OS countermeasures | None tracked | None shipped | Stable | Zoom, Google Workspace, Webex, macOS and Windows all shipped nothing in-quarter that detects or blocks a capture-excluded overlay |
| CORRECTION — open-source population | We reported 20+ repositories and 8+ forks | 10 active projects at 50+ stars, and none of them is a fork | ↓ Our figure was wrong | The Q2 count was not reproducible under any stated bar, and every active project is an independent reimplementation rather than a fork |
| CORRECTION — Parakeet AI pricing | We reported $20–40/month | $149.90/month on the vendor's own page | ↓ Our figure was wrong | Our figure appears to have come from a promotional price carried by third-party review sites. It understated the real cost by 4–7× |
| CORRECTION — Final Round AI capability | We listed audio earpiece integration for behavioural rounds | Contradicted by the vendor: “silent — it never speaks for you” | ↓ Withdrawn | We could not substantiate our own claim against any vendor source and have removed it |
§9
Scoring the five forecasts we made for Q3.
The Q2 edition of this index made five specific predictions about Q3. Two were right, three were wrong, and one of the wrong ones was wrong in a way that taught us more than the correct ones did. We publish the scoring because a quarterly index that never grades itself is marketing.
| What we predicted | Verdict | What actually happened |
|---|---|---|
| Cluely ships hardware (earpieces, smart glasses); hardware category rises 75 → 85+ | WRONG | No hardware product shipped. We found no Cluely product announcement in the quarter at all; the company's own blog has published nothing since 9 June 2026. We have moved the hardware category down 2 points rather than up 10 |
| Compact models cross the threshold; on-device mobile LLMs emerge as a documented vector | CORRECT | And understated. A 2.6B model now ships in a 1.59 GB file that its vendor reports running at 30 tokens/second on a phone, while a 30B multimodal model became an Apache-2.0 download at 16.76 GB |
| Cross-session intelligence becomes standard among assessment platforms; proxy scores fall | WRONG, twice | No assessment platform shipped cross-session detection in the window. Our stated premise was also unsupported: CodeSignal's Suspicion Score is session-isolated and never claimed otherwise. The capability did ship — from Sardine, a financial-crime fraud vendor selling into applicant tracking systems, in May 2026. We were watching the wrong layer of the stack |
| EU AI Act employment provisions take effect in stages through 2026; vendors face exposure | WRONG, and backwards | The opposite happened. Regulation (EU) 2026/1744, in force 27 July 2026, deferred the Annex III high-risk obligations covering recruitment from 2 August 2026 to 2 December 2027. Separately, our reasoning was flawed: the US state laws we pointed to regulate employers and hiring-AI vendors, not the candidate-side tools in this index |
| Agent-mode cheating: first movers expected Q3–Q4 | CORRECT | No commercial candidate-side product has shipped autonomous capability. We checked thirteen vendors' own product pages; every one remains query-response with the candidate as the actuator. Engineering effort went into undetectability, latency and off-screen delivery instead |
What we are watching for Q4 2026
- Whether remote-assist becomes a category. One vendor now sells hands-on remote control. If a second ships the same feature, this stops being a product decision and becomes the category's direction of travel.
- Whether macOS capture exclusion degrades further. ScreenCaptureKit already breaks the guarantee these tools rely on. If Apple extends that behaviour, the macOS half of this threat class weakens without any proctoring vendor doing anything.
- Whether the offline path reaches non-technical candidates. Local inference currently requires installing a runtime. The moment one commercial overlay bundles a local model as a default setting, the network-visible population of this threat drops sharply.
- Whether capture-exclusion detection reaches production. The primitive is public, and the ground is no longer unclaimed: a startup called InterviewWatch launched on 8 July 2026 specifically to detect AI assistance during live technical interviews. No established proctoring vendor has shipped the capability. The question for Q4 is whether any of them buys or builds it, or whether detection of this threat class stays with new entrants.
Methodology & citation
How scores are calculated. How to cite this index.
Scoring methodology
Threat score (1–10 per tool, 1–100 per category): Weighted composite of: (a) detection difficulty against deployed defenses, 40%; (b) ease of use for the median attacker, 30%; (c) prevalence in the wild, sourced from platform reports and researcher observation, 20%; (d) capability ceiling ( i.e., maximum sophistication achievable), 10%.
Composite index score: Category scores weighted by observed usage distribution across real assessment events. Commercial overlay weight: 25%. Open-source projects: 25%. On-device LLMs: 20%. Remote access: 10%. Proxy services: 10%. Hardware: 10%. Browser automation is reported for continuity but carries no weight in the composite, because we no longer observe it as a primary vector. Applying these weights to the Q3 category scores gives 89.40, reported as 89.
Detection verdicts: BYPASSED = documented, working evasion technique accessible to a motivated candidate with ≤8 hours of preparation. PARTIALLY BYPASSED = evasion exists but requires significant preparation or has meaningful false-positive costs. NO KNOWN BYPASS = no documented evasion technique for the stated scope.
Update cadence: Quarterly. This Q3 2026 edition covers July through early September 2026, with all vendor pricing and repository data collected through 2 September 2026. Next update: December 2026 (Q4).
APA 7th citation
Bhanushali, D. (2026, September). AI Cheating Threat Index: Q3 2026. Aiseptor. https://aiseptor.com/research/threat-index
Data sources
- Aiseptor Threat Intelligence: reverse-engineered tool analysis
- GitHub repositories: OpenCluely, Pluely, Natively, MindWhisperAI, ShadeCoder, LeetcodeWizard
- CodeSignal Fraud Rate Report (Feb 2026)
- Fabric: 19,368 AI interview analysis (Jan 2026)
- Talview AI Threat Index Report 2026
- TechCrunch: overlay tool reporting (2025–2026)
- arXiv 2601.17280: keystroke dynamics study (Jan 2026)
- FBI IC3 advisory: state-sponsored hiring fraud
- Experian 2026 Fraud Forecast
- Gartner: fabricated profile projections
Frequently Asked Questions
What is the composite threat score in the Q3 2026 index?
89 out of 100, up from 87 in Q2. The two largest movers were on-device LLMs (+6), after a 30-billion-parameter multimodal model was released under Apache-2.0 in a 16.8 GB file that runs on a consumer laptop, and remote-access tools (+7), after a commercial overlay vendor began selling live remote control of the candidate's machine as a product feature.
How many exam-security defenses have a documented bypass?
6 of the 7 defense approaches tracked in this index have a documented, working bypass against current AI cheating tools. That did not change in Q3: we checked the release notes of Zoom, Google Workspace, Webex, macOS and Windows across the quarter and found no vendor shipped a capability that detects or blocks a screen-capture-excluded overlay. Only network-layer enforcement, which removes the network conditions these tools require rather than trying to detect them by name, has no known bypass in this quarter's tracking.
How often is the Threat Index updated?
Quarterly. This edition covers Q3 2026, with data collected through 2 September 2026, and tracks quarter-over-quarter deltas against the Q2 2026 baseline across commercial overlays, open-source projects, on-device LLMs, remote-access tools, and proxy/exam-fraud services.
Related research
Continue reading
Annual report
AI Cheating Statistics 2026
The comprehensive annual dataset: fraud rates, candidate attitudes, regional breakdowns, bad-hire cost, and the full bypass map.
Read →Research hub
All research
Primary-source research on AI-assisted cheating, assessment fraud, and network-layer prevention architectures.
Read →Reference
Glossary
Definitions of the tools, techniques, and architectural terms used throughout our research reports.
Read →The only rated defense with no known bypass
One defense passed the rating. See how it works live.
Every tool in this index (commercial overlays, open-source forks with custom binary names, remote-access pipelines) requires a network path to function. Aiseptor closes that path at the OS level before the assessment begins. Book a 30-minute demo and watch us block the Q3 2026 toolkit in real time on a candidate device.