Deoptimization: falling off the cliff
Deopt abandons optimised machine code and resumes in Ignition at the exact bytecode offset. Eager vs lazy deopt, the frame-translation mechanism that reconstructs the interpreter frame, the catastrophic deopt-loop, and the triggers and flags to diagnose it.
A function that ran at 40 ns per call in your benchmark suddenly costs 4 µs per call in production — a hundred times slower than uncompiled would have been. There is no infinite loop, no blocking I/O, no GC storm. What you are watching is a function being optimised and thrown away, optimised and thrown away, dozens of times a second, because one of TurboFan’s guards keeps failing on a value it can’t stop seeing. This is the deopt-loop, and it is the single most expensive way to be wrong about types in JavaScript.
What deopt actually is
A guard failure (lesson 04) does not crash and does not silently produce a wrong answer. It triggers a deoptimization: V8 abandons the optimised machine code for that function and resumes execution in the Ignition interpreter at the exact bytecode offset corresponding to where the optimised code gave up. Execution continues correctly, just slower. The optimised code is discarded (or marked invalid), and the function falls back to its lower tier; if it stays hot and its feedback stabilises, it can be re-optimised later.
There are two flavours, distinguished by what invalidated the assumption.
Eager deopt happens synchronously when a guard fails at runtime — the function is executing optimised code, hits a CheckMaps / CheckSmi / CheckBounds, the check fails, and control bails out right there. The trigger is local: this call saw a value the speculation didn’t allow. In --trace-deopt it appears as DEOPT eager with a reason like wrong map or not a Smi.
Lazy deopt happens when an assumption is invalidated externally, not by the function’s own execution. You redefine a function the optimised code inlined, mutate a prototype it depended on, add a property to a shared object, or change a global. V8 cannot reach into a frame that may be running, so instead it marks all dependent optimised code for deopt on next entry — the code is “lazily” deoptimised the next time it would be called. In --trace-deopt it appears as DEOPT lazy. The classic cause is monkey-patching: Array.prototype.map = ... after hot code has been optimised against the original.
Frame translation: how it returns cleanly
When the guard fails, V8 must resume in the interpreter — but the two frames speak entirely different languages. The hard part of deopt is that the optimised frame and the interpreter frame are completely different layouts. Optimised code keeps values in machine registers and a packed stack frame, often in unboxed representations (a raw int32, a raw float64). The interpreter expects values in its own register file (the interpreter’s “registers” are stack slots) in their tagged, boxed form. To resume in the interpreter, V8 must rebuild the interpreter’s frame from the optimised one.
It does this with deopt metadata recorded at compile time: for each possible deopt point, TurboFan stores a description mapping every live optimised-frame location (this machine register holds local i, this stack slot holds total, this is an unboxed double that must be re-boxed) to the interpreter register it belongs in. At deopt time the deoptimizer reads that metadata and performs frame translation — it allocates and fills a fresh interpreter frame with every live value put back where Ignition expects it, re-boxing unboxed values, then jumps into the interpreter at the saved bytecode offset. A single deopt’s translation is cheap, on the order of microseconds.
The deopt-loop: the actual catastrophe
One deopt is fine — microseconds of bookkeeping, then the function re-warms with feedback that now includes the surprising case, re-optimises with a broader assumption, and stays optimised. The catastrophe is the deopt-loop: a function is repeatedly fed a shape or type it cannot keep stable, so it optimises, deopts, re-optimises (the function is still hot, so V8 promotes it again), deopts again — forever. Each cycle pays the deopt translation plus a fresh TurboFan compile (tens to hundreds of ms) plus the time spent running in the slow tier in between. A function in a deopt-loop is orders of magnitude slower than if you had simply disabled optimisation with --no-opt.
When you see a function oscillating between optimised and deoptimised state in —trace-deopt, look for one of these recurring patterns — each is a way to make a guard fail on a recurring value:
- Smi overflow — an arithmetic site optimised on
SignedSmallproduces a value past the Smi range (greater than 2^30-ish), failingCheckSmievery time the big value recurs. - Shape change / hidden-class divergence — a call site sees objects with diverging maps;
CheckMapsfails. Adding properties in different orders, conditional fields,delete. - Changing a field’s type — a field optimised as Smi later holds a string or object, so loads guarded on the field’s representation deopt.
- Reading
argumentsin optimised functions — historically forced deopt or blocked optimisation; modern V8 handles many cases but rest parameters are still the safe choice. - Out-of-bounds or holey access —
CheckBoundsfailing, or transitioning a packed array to holey elements kind.
You diagnose it with --trace-deopt (every deopt with its function, bytecode offset, and reason) and, in d8 under --allow-natives-syntax, %GetOptimizationStatus(fn) to read the function’s tier bits. The browser-track lesson TurboFan’s speculative engine and the deopt-loop trap walks a full trace-and-fix; here the emphasis is the mechanism and the fix pattern: route the pathological value to a separate slow path before the hot block so the fast path stays type-stable.
- Single deopt translation
- ~microseconds
- Re-optimise (TurboFan)
- tens-hundreds of ms
- Deopt-loop slowdown
- orders of magnitude
- Eager deopt trigger
- guard fails at runtime
- Lazy deopt trigger
- external invalidation
- Smi range
- ~31-bit signed
- Diagnose flag
- --trace-deopt
- Status intrinsic
- %GetOptimizationStatus(fn)
You monkey-patch `Array.prototype.includes` at runtime, after a hot function that used it has been TurboFan-optimised. What kind of deopt does this cause?
Why is a deopt-loop slower than running the same function with `--no-opt` (optimisation disabled entirely)?
Order what happens during a single eager deopt.
- 1 Optimised code executes and reaches a guard (e.g. CheckSmi)
- 2 The guard fails: the value violates the speculated type
- 3 The deoptimizer reads the compile-time deopt metadata
- 4 It rebuilds the interpreter frame, re-boxing live values into interpreter registers
- 5 Execution resumes in Ignition at the saved bytecode offset
▸Common mistake
A subtle deopt-loop hides in “defensive” code that returns mixed types: a parser that returns a number for valid input but null for invalid, feeding a hot arithmetic site. For 99% of inputs the site is Smi-stable and optimises; the occasional null fails CheckSmi and deopts; the site re-optimises on Smi again because nulls are rare; the next null deopts again. The fix is not “handle null faster” — it is to split the stream: validate and discard nulls before the hot loop so the arithmetic site only ever sees numbers.
- 01Distinguish eager and lazy deopt with a trigger for each.
- 02Explain frame translation and why it is needed.
- 03What is a deopt-loop, why is it catastrophic, and how do you fix it?
Deoptimization is the safety valve that keeps speculative optimisation correct. When an assumption breaks, V8 abandons the optimised machine code and resumes in the Ignition interpreter at the exact bytecode offset where the optimised code gave up — execution stays correct, just slower. Eager deopt is synchronous: the function’s own optimised code hits a failing guard (CheckMaps ‘wrong map’, CheckSmi ‘not a Smi’, CheckBounds ‘out of bounds’). Lazy deopt is external: redefining an inlined function, mutating a prototype, or changing a global invalidates dependent optimised code, which V8 marks for deopt on next entry because it cannot patch a running frame. The clean return is possible thanks to frame translation: TurboFan records deopt metadata mapping every live optimised-frame value (in registers, often unboxed) to the interpreter register it belongs in, and at deopt time the deoptimizer rebuilds a tagged interpreter frame and jumps to the saved offset. A single deopt costs microseconds and is the system working as designed. The catastrophe is the deopt-loop: a recurring value breaks a guard on essentially every call, so V8 optimises and deopts forever, paying a fresh compile each cycle and ending up orders of magnitude slower than —no-opt. Triggers include Smi overflow, shape/hidden-class divergence, changing a field’s type, reading arguments, and holey/out-of-bounds access. Diagnose with —trace-deopt and %GetOptimizationStatus; fix by routing the pathological value to a slow path before the hot block so the fast path stays type-stable. Now when you see a 100× production regression with no obvious cause, open —trace-deopt before reaching for a profiler — a deopt-loop will announce itself immediately.
Practice
Start at the top. Tasks go easiest → hardest: recall a fact, apply it to a case, then a senior-level stretch. Open one, attempt it, then reveal.
appears again in184
- Why GraphQL gets N+1junior
- DataLoader mechanics: tick-boundary batchingmiddle
- Batch function contracts: ordering, shapes, errorsmiddle
- Federation and lookahead: batching beyond DataLoadermiddle
- Query complexity defences: depth, cost, persisted queriesmiddle
- Senior GraphQL API: scheduling contract, tenant isolation, observabilitysenior
- Why idempotency: making retries safejunior
- Server-side state machine: four states of an idempotency keymiddle
- Outbox and inbox: effectively-once across the dual-write boundarymiddle
- Concurrency and cache architecture for idempotency at scalesenior
- Observability, production failures, and global-scale designsenior
- The event loop: one thread, three queuesjunior
- Tasks, microtasks, and scheduler.yield()middle
- Microtask starvation, Long Tasks, and LoAFsenior
- Node.js event loop: phases, nextTick, and loop lagsenior
- React, Vue, and INP observability in productionsenior
- The render pipeline: six stages from bytes to pixelsjunior
- Stage costs and the renderer process modelmiddle
- Invalidation, dirty bits, and containmiddle
- Compositor layers: promotion, overlap, and GPU memorymiddle
- DevTools flame strip and the frame lifecyclemiddle
- Layout thrash: forced synchronous layoutsenior
- BeginMainFrame, compositor-driven animations, and GPU memorysenior
- Production observability: LoAF, INP, and the full attack surfacesenior
- What V8 is and why performance varies 100×junior
- V8''''s four-tier JIT pipeline and profile-guided tieringmiddle
- Hidden classes, transition trees, and memory layoutmiddle
- Inline caches, IC states, and deoptimizationmiddle
- Orinoco GC: parallel scavenger, concurrent marking, and write barriersmiddle
- TurboFan''''s speculative engine and the deopt-loop trapsenior
- V8 in production: isolates, pointer compression, and real failuressenior
- Service worker lifecycle and cache strategiesmiddle
- Service worker edge cases: version skew, durability, and navigation trapssenior
- What the reconciler does: render vs commitjunior
- The fiber object and the double-buffer treemiddle
- Render phase purity and commit phase sub-stepsmiddle
- Reconciliation: diffing heuristics and the key trapmiddle
- Priority lanes, time-slicing, and useTransitionmiddle
- Bailout, memoisation, and tearingsenior
- React Profiler, the Compiler, and production observabilitysenior
- Rendering strategies: SSG, SSR, ISR, streaming, and hydrationjunior
- SSG, SSR, ISR, streaming, and RSC — how each worksmiddle
- Hydration cost: selective, progressive, islands, resumabilitymiddle
- Hydration mismatch: causes, detection, and the determinism rulesenior
- RSC, per-route strategy, and production observabilitysenior
- Core Web Vitals: what LCP, INP, and CLS measurejunior
- CLS: why layout shifts happen and how to stop themmiddle
- Metric tradeoffs, RUM attribution, and the CI+field loopsenior
- The full picture: URL to LCP to INP as a relay racejunior
- Eight layers traced: from the service worker to the second navigationmiddle
- Five canonical breaks: where production reliably diessenior
- The three-track method: reading traces and building a monitored systemsenior
- What is a cache stampede and why it makes things worsejunior
- Lock and single-flight: bounding concurrent rebuildsmiddle
- XFetch: coordination-free probabilistic early expirationmiddle
- Stale-while-revalidate and CDN request coalescingmiddle
- Detecting stampedes and designing TTL for productionmiddle
- Metastable failure, fencing tokens, and production postmortemssenior
- What a relation is: tables, rows, keys, and constraintsjunior
- Constraints, keys, and Postgres data typesmiddle
- Normal forms, denormalization, and why schemas stickmiddle
- JSONB, arrays, and when a side table winsmiddle
- Heap storage, TOAST, and column alignmentsenior
- Schema integrity: deferral, versioning, and production failure modessenior
- Relational vs document, wide-column, graph, and key-valuesenior
- Index-only scans, the Visibility Map, and INCLUDEsenior
- Production failure modes and the index audit playbooksenior
- pg_statistic, ANALYZE, and production observabilitymiddle
- Production failure modes and plan stabilitysenior
- MVCC: why readers and writers never wait for each otherjunior
- Row versions and snapshots: the on-disk mechanicsmiddle
- HOT updates and isolation levels: what you gain and what you paymiddle
- Vacuum and bloat: keeping the storage tax boundedmiddle
- CLOG, XID wraparound, and MultiXact: deep visibility internalssenior
- SSI internals and production autovacuum tuningsenior
- Real-world MVCC failures, deployment patterns, and distributed snapshotssenior
- Connection pools: amortising the cost of a Postgres backendjunior
- PgBouncer session, transaction, and statement modesmiddle
- Pool sizing: the (cores × 2) + spindles formula and the two-layer stackmiddle
- Pool exhaustion and idle-in-transaction: the 3 AM failure modemiddle
- Migrating to transaction mode: rollout playbook and PgBouncer 1.21 prepared statementsmiddle
- The Postgres process model and why raising max_connections degrades throughputsenior
- Pooler landscape 2026, serverless connection storms, and the full failure-mode taxonomysenior
- What a schema migration is and why it replaces ad-hoc DDLjunior
- ADD COLUMN: instant in PG 11+ vs rewrite in older Postgresjunior
- The lock-queue failure mode: why instant DDL can freeze the databasemiddle
- Safe DDL patterns: NOT VALID, CONCURRENTLY, and unsafe-op fixesmiddle
- Expand-contract: zero-downtime for breaking schema changesmiddle
- Advisory locks, migration tools, and deploy coordinationsenior
- Migration failure taxonomy and production disciplinesenior
- Why sharding exists: the single-Postgres ceilingjunior
- Shard-key selection: hash, range, list, and directory strategiesmiddle
- Partitioning vs sharding: same word, two different thingsmiddle
- Co-location and Citus: the invariant that makes sharding usablemiddle
- The hot-shard failure mode: detection, isolation, and durable policymiddle
- Schema-based sharding and multi-tenancy alternativessenior
- Online resharding, 2PC, and the operational cost of shardingsenior
- The seven acts: from CREATE TABLE to Citusjunior
- Acts 1–3 in depth: schema, indexes, and planner statisticsmiddle
- Acts 4–6 in depth: MVCC bloat, connection pooling, and safe migrationsmiddle
- Act 7 in depth: sharding, co-location, and the seven-tier tradeoff cascademiddle
- Observability, anti-patterns, and production triagesenior
- Raft roles, terms, and why majority quorums prevent split brainjunior
- How Raft replicates a log entry and decides it is safe to commitmiddle
- Raft leader election: timeouts, voting rules, and the four safety propertiesmiddle
- Raft in the real world: partitions, slow disks, and client routingmiddle
- Raft extensions: pre-vote, learners, snapshots, and linearizable readssenior
- Raft in production: membership changes, Multi-Raft, and observabilitysenior
- Where data fetching happens — and why it decides LCPjunior
- Fetch waterfalls — diagnosis and the Promise.all curemiddle
- React Server Components and Suspense streamingmiddle
- Client-side cache: TanStack Query, SWR, and stale-while-revalidatemiddle
- LCP, prefetch, and race conditions in interactive fetchingmiddle
- Senior internals: RSC payload, caching layers, and production failure modessenior
- The three-way handshakejunior
- Sequence numbers and connection statemiddle
- DNS: what it does and why it existsjunior
- The resolver walk: referrals, record types, and gluemiddle
- TTL, caching, and DNS propagationmiddle
- The 1-RTT handshake: key shares and ECDHEmiddle
- Session resumption and 0-RTTmiddle
- WebSocket: the HTTP upgrade handshakejunior
- WebSocket frame format: opcodes, masking, fragmentationmiddle
- WebSocket backpressure: when clients can''''t keep upmiddle
- Reconnection: jittered backoff, thundering herd, message resumptionsenior
- WebSocket at scale: HTTP/2 multiplexing, permessage-deflate, C10Msenior
- WebSocket in production: proxies, security, and distributed architecturesenior
- What reverse proxies dojunior
- Health checks, connection draining, and slow startmiddle
- Session affinity, consistent hashing, and the right fixmiddle
- Retry storms, circuit breakers, and load sheddingsenior
- Resilient LB architecture: anycast, zone-aware routing, and observabilitysenior
- Why QUIC and not TCP+TLSjunior
- Connection IDs and network migrationmiddle
- 0-RTT resumption and packet encryptionsenior
- DDoS: what it is and why it worksjunior
- Amplification attacks and state exhaustionmiddle
- Rate limiting: algorithms and architecturemiddle
- WAFs, firewalls, mTLS, and HSTSmiddle
- DNS cache poisoning and BGP hijackingsenior
- Defense-in-depth architecture and attack economicssenior
- DNS, TCP, TLS in sequence: where the milliseconds gomiddle
- Proxy intercepts and security gates: rate limiters, WAF, mTLSmiddle
- Alternate paths: QUIC 0-RTT, WebSocket upgrade, connection migrationmiddle
- Observability: distributed traces, USE/RED, and samplingsenior
- Resilience: cascading retries, circuit breakers, and error budgetssenior
- What the three signals are: logs, metrics, and tracesjunior
- Why structured logs exist: the diary vs the spreadsheetjunior
- The production log schema: fields every line must carrymiddle
- PII redaction and log injectionsenior
- OTel Logs Data Model and audit logs as a subsystemsenior
- SLI, SLO, and the error budget: reliability by the numbersjunior
- Error budget policy, latency SLOs, and composite journeysmiddle
- Production SLO failures, self-observability, security, and the big picturesenior
- The incident loop: from pager to postmortem to preventionmiddle
- Cache lines, struct layout, and false sharingmiddle
- SIMD, SoA vs AoS, and memory bandwidthmiddle
- Cache-oblivious algorithms, PGO, and production failuressenior
- GC in production: observability, security, edge cases, and fleet governancesenior
- Batching: amortize fixed cost per operationjunior
- The batching window: size and wait timemiddle
- Batching in Kafka and Postgresmiddle
- io_uring and observability of batchingmiddle
- From Nagle to io_uring: evolution of batchingmiddle
- Backpressure, failure isolation, and batch security in productionsenior
- CI enforcement and RUM: making budgets stickmiddle
- V8 JIT pipeline, HTTP priorities, and bundle securitysenior
- The performance loop: discipline, not a projectjunior
- Classify and fix: matching bottleneck families to remediesmiddle
- Observability stack and CI gates: catching regressions before they shipmiddle
- Incident to enforcement: SLO burn to verified fix in 35 minutesmiddle
- Culture, economics, and org-scale performancesenior
- At-most-once, at-least-once, exactly-once: the three delivery contractsjunior
- The three failure legs — where duplicates and losses actually happenmiddle
- Consumer-side dedup: the cheapest path to exactly-once processingmiddle
- Kafka exactly-once semantics: idempotent producer and transactionsmiddle
- SQS visibility timeout, DLQ, and the outbox patternmiddle
- Exactly-once in production: impossibility proof, hybrid patterns, and real incidentssenior
- What OAuth is and why passwords are not the answerjunior
- Authorization code flow with PKCEmiddle
- ID token validation and JWKS cache managementmiddle
- Refresh token rotation and scope-based least privilegemiddle
- Sender-constrained tokens: DPoP and mTLSsenior
- OAuth in production: audience attacks, observability, and real failuressenior
Something unclear?
Ask a question about this lesson. Questions are anonymous and go straight to the author to make the lesson better.
Apply this
Put this lesson to work on a real build.