On-Stack Replacement: swapping the frame mid-loop
A function called once but containing a long hot loop never returns to be re-entered, so normal tier-up can't speed it up. OSR swaps the executing frame in flight
You write a script with one giant for loop that crunches ten million rows. The function is called exactly once — main() — and never returns until the loop is done. By the tiering rules you have learned, this function can never speed up: tier-up swaps the optimised code in on the next call, and there is no next call. Yet you measure it and the loop does accelerate, partway through, while it is still running. The only way that can happen is if V8 replaced the code under the running loop’s feet. That trick is On-Stack Replacement.
The problem normal tier-up can’t solve
Recall how tier-up installs optimised code (lesson 01): when a function’s budget trips, V8 compiles a faster version and swaps it in for the next entry to the function. That model has a blind spot. Consider a function that is called once and spends all its time in a single long loop:
function crunch(rows) {
let acc = 0;
for (let i = 0; i < rows.length; i++) { // 10,000,000 iterations
acc += score(rows[i]);
}
return acc;
}
crunch(tenMillionRows); // called exactly onceThe loop is blazingly hot — millions of back-edges — so its back-edge counter trips the budget almost immediately (lesson 01 introduced this counter precisely for this case). V8 wants to optimise. But the normal mechanism is useless here: it would install optimised code for the next call to crunch, and crunch is mid-flight on the only call it will ever get. By the time crunch returns, the work is done. Waiting for re-entry means never optimising the one loop that matters.
OSR: replacing the frame in flight
On-Stack Replacement is the mechanism that breaks this deadlock. Instead of waiting for the function to return and be re-called, V8 replaces the executing frame in place, mid-loop. The sequence:
- The loop’s back-edge counter trips the tier-up budget while the loop is running in the interpreter (or baseline tier).
- V8 compiles a special OSR-entry version of the function — optimised machine code whose entry point is not the function’s top, but the loop header. It is specialised to begin executing as if control had just arrived at the start of a loop iteration, with all the loop’s live variables already in place.
- V8 performs a live-state transfer: it reads the current values of every live variable from the interpreter frame (the loop index
i, the accumulatoracc, therowsreference) and writes them into the optimised frame’s registers and stack slots, in the representations the optimised code expects. This is the inverse of the deopt frame translation from lesson 05 — there we rebuilt an interpreter frame from an optimised one; here we build an optimised frame from an interpreter one. - Control transfers into the OSR-entry optimised code at the loop header, and the same loop continues — but now in optimised machine code. The remaining nine-and-a-half million iterations run fast.
- After the loop finishes and
cruncheventually returns, normal optimised code (a regular, top-entry compilation) takes over for any subsequent calls.
Why micro-benchmarks live or die by OSR
This is the mechanism behind a notorious benchmarking pitfall. If your benchmark fires up once and runs a single giant loop, ask yourself: has your function ever been called enough times to get a normal tier-up? The answer is almost certainly no — you are depending entirely on OSR. A micro-benchmark that wraps its work in a single huge loop — for (let i = 0; i < 1e9; i++) work() — does not warm up work through invocation-count tier-up in the usual way; it relies on OSR kicking in partway through to optimise the loop body live. That has two consequences for anyone measuring performance:
- The first chunk of iterations runs unoptimised (interpreter/baseline), then there is a discontinuity where OSR fires and the rest run fast. If you time the whole loop as one number, you blend cold and hot execution and get a meaningless average.
- OSR-entry code can be slightly worse than a normal top-entry optimisation, because it must accept the loop’s live state as it found it rather than from a clean function entry, which constrains some optimisations. So an OSR-warmed measurement can under-report the steady-state speed you would see if the function were called many times normally.
The fix is the standard one: warm up the function with many separate calls before timing (so it gets a normal, non-OSR optimisation), and measure steady-state iterations, not the warm-up transient. You can see OSR happen with --trace-osr; combined with --trace-opt it shows the OSR-entry compile distinct from the regular one.
- Triggered by
- loop back-edge counter
- Entry point
- loop header (not fn top)
- State moved
- interpreter frame -> optimised
- Relation to deopt
- inverse frame translation
- OSR code vs normal opt
- sometimes slightly worse
- After the loop
- normal opt for next calls
- Trace flag
- --trace-osr
- Benchmark relevance
- single-giant-loop benches
Why can't normal tier-up optimise a function that is called once and spends all its time in one long loop?
What does OSR have to do at the moment it swaps the frame, and how does it relate to deopt?
Order how OSR optimises a long-running loop in a once-called function.
- 1 The loop's back-edge counter trips the tier-up budget mid-execution
- 2 V8 compiles an OSR-entry version whose entry point is the loop header
- 3 V8 transfers live loop state from the interpreter frame into the optimised frame
- 4 Control resumes in the optimised code and the same loop continues fast
- 5 After the loop returns, normal optimised code serves any later calls
▸Why this works
Why does the OSR-entry version sometimes optimise slightly worse than a normal compilation? Because a normal top-entry optimisation gets to assume a clean function entry — fresh arguments, no live loop state mid-computation — and can lay out the whole function optimally. The OSR-entry version must instead accept the loop’s live values exactly as they exist at the trip point and begin from the loop header, which pins some representations and limits a few rearrangements the compiler would otherwise make. It is still vastly faster than the interpreter; it is just occasionally a hair behind the steady-state code you get from many normal calls — which is exactly why warming with separate calls gives cleaner benchmark numbers.
- 01What problem does OSR solve that normal tier-up cannot?
- 02Walk through the OSR mechanism step by step.
- 03Why must single-giant-loop micro-benchmarks be warmed with separate calls, and what does OSR have to do with it?
On-Stack Replacement is the mechanism that lets V8 optimise a loop that is already running, closing the one gap in normal tier-up. Normal tier-up installs optimised code for the next entry to a function, which is useless for a function called once that spends all its time in a single long loop: the loop is intensely hot via its back-edge counter, but there is no next call to install code for. OSR fixes this by replacing the executing frame in flight. When the back-edge budget trips, V8 compiles an OSR-entry version of the function whose entry point is the loop header, then performs a live-state transfer - reading every live loop variable out of the interpreter frame and writing it into the optimised frame in the representations the optimised code expects, the exact inverse of deopt’s frame translation. Control jumps into the optimised code at the loop header and the same in-flight loop continues fast, with normal top-entry optimised code taking over for any later calls once the loop returns. Because single-giant-loop micro-benchmarks rely on OSR rather than ordinary tier-up, and because OSR-entry code can be marginally worse than a clean optimisation, such benchmarks must be warmed with many separate calls and measured at steady state. Observe it all with —trace-osr alongside —trace-opt. Now when you write a benchmark and wonder why the first million iterations are slow and the rest fast, you know: that is the OSR boundary, and the fix is to call your function in a warm-up loop before you start the clock.
Practice
Start at the top. Tasks go easiest → hardest: recall a fact, apply it to a case, then a senior-level stretch. Open one, attempt it, then reveal.
appears again in184
- Why GraphQL gets N+1junior
- DataLoader mechanics: tick-boundary batchingmiddle
- Batch function contracts: ordering, shapes, errorsmiddle
- Federation and lookahead: batching beyond DataLoadermiddle
- Query complexity defences: depth, cost, persisted queriesmiddle
- Senior GraphQL API: scheduling contract, tenant isolation, observabilitysenior
- Why idempotency: making retries safejunior
- Server-side state machine: four states of an idempotency keymiddle
- Outbox and inbox: effectively-once across the dual-write boundarymiddle
- Concurrency and cache architecture for idempotency at scalesenior
- Observability, production failures, and global-scale designsenior
- The event loop: one thread, three queuesjunior
- Tasks, microtasks, and scheduler.yield()middle
- Microtask starvation, Long Tasks, and LoAFsenior
- Node.js event loop: phases, nextTick, and loop lagsenior
- React, Vue, and INP observability in productionsenior
- The render pipeline: six stages from bytes to pixelsjunior
- Stage costs and the renderer process modelmiddle
- Invalidation, dirty bits, and containmiddle
- Compositor layers: promotion, overlap, and GPU memorymiddle
- DevTools flame strip and the frame lifecyclemiddle
- Layout thrash: forced synchronous layoutsenior
- BeginMainFrame, compositor-driven animations, and GPU memorysenior
- Production observability: LoAF, INP, and the full attack surfacesenior
- What V8 is and why performance varies 100×junior
- V8''''s four-tier JIT pipeline and profile-guided tieringmiddle
- Hidden classes, transition trees, and memory layoutmiddle
- Inline caches, IC states, and deoptimizationmiddle
- Orinoco GC: parallel scavenger, concurrent marking, and write barriersmiddle
- TurboFan''''s speculative engine and the deopt-loop trapsenior
- V8 in production: isolates, pointer compression, and real failuressenior
- Service worker lifecycle and cache strategiesmiddle
- Service worker edge cases: version skew, durability, and navigation trapssenior
- What the reconciler does: render vs commitjunior
- The fiber object and the double-buffer treemiddle
- Render phase purity and commit phase sub-stepsmiddle
- Reconciliation: diffing heuristics and the key trapmiddle
- Priority lanes, time-slicing, and useTransitionmiddle
- Bailout, memoisation, and tearingsenior
- React Profiler, the Compiler, and production observabilitysenior
- Rendering strategies: SSG, SSR, ISR, streaming, and hydrationjunior
- SSG, SSR, ISR, streaming, and RSC — how each worksmiddle
- Hydration cost: selective, progressive, islands, resumabilitymiddle
- Hydration mismatch: causes, detection, and the determinism rulesenior
- RSC, per-route strategy, and production observabilitysenior
- Core Web Vitals: what LCP, INP, and CLS measurejunior
- CLS: why layout shifts happen and how to stop themmiddle
- Metric tradeoffs, RUM attribution, and the CI+field loopsenior
- The full picture: URL to LCP to INP as a relay racejunior
- Eight layers traced: from the service worker to the second navigationmiddle
- Five canonical breaks: where production reliably diessenior
- The three-track method: reading traces and building a monitored systemsenior
- What is a cache stampede and why it makes things worsejunior
- Lock and single-flight: bounding concurrent rebuildsmiddle
- XFetch: coordination-free probabilistic early expirationmiddle
- Stale-while-revalidate and CDN request coalescingmiddle
- Detecting stampedes and designing TTL for productionmiddle
- Metastable failure, fencing tokens, and production postmortemssenior
- What a relation is: tables, rows, keys, and constraintsjunior
- Constraints, keys, and Postgres data typesmiddle
- Normal forms, denormalization, and why schemas stickmiddle
- JSONB, arrays, and when a side table winsmiddle
- Heap storage, TOAST, and column alignmentsenior
- Schema integrity: deferral, versioning, and production failure modessenior
- Relational vs document, wide-column, graph, and key-valuesenior
- Index-only scans, the Visibility Map, and INCLUDEsenior
- Production failure modes and the index audit playbooksenior
- pg_statistic, ANALYZE, and production observabilitymiddle
- Production failure modes and plan stabilitysenior
- MVCC: why readers and writers never wait for each otherjunior
- Row versions and snapshots: the on-disk mechanicsmiddle
- HOT updates and isolation levels: what you gain and what you paymiddle
- Vacuum and bloat: keeping the storage tax boundedmiddle
- CLOG, XID wraparound, and MultiXact: deep visibility internalssenior
- SSI internals and production autovacuum tuningsenior
- Real-world MVCC failures, deployment patterns, and distributed snapshotssenior
- Connection pools: amortising the cost of a Postgres backendjunior
- PgBouncer session, transaction, and statement modesmiddle
- Pool sizing: the (cores × 2) + spindles formula and the two-layer stackmiddle
- Pool exhaustion and idle-in-transaction: the 3 AM failure modemiddle
- Migrating to transaction mode: rollout playbook and PgBouncer 1.21 prepared statementsmiddle
- The Postgres process model and why raising max_connections degrades throughputsenior
- Pooler landscape 2026, serverless connection storms, and the full failure-mode taxonomysenior
- What a schema migration is and why it replaces ad-hoc DDLjunior
- ADD COLUMN: instant in PG 11+ vs rewrite in older Postgresjunior
- The lock-queue failure mode: why instant DDL can freeze the databasemiddle
- Safe DDL patterns: NOT VALID, CONCURRENTLY, and unsafe-op fixesmiddle
- Expand-contract: zero-downtime for breaking schema changesmiddle
- Advisory locks, migration tools, and deploy coordinationsenior
- Migration failure taxonomy and production disciplinesenior
- Why sharding exists: the single-Postgres ceilingjunior
- Shard-key selection: hash, range, list, and directory strategiesmiddle
- Partitioning vs sharding: same word, two different thingsmiddle
- Co-location and Citus: the invariant that makes sharding usablemiddle
- The hot-shard failure mode: detection, isolation, and durable policymiddle
- Schema-based sharding and multi-tenancy alternativessenior
- Online resharding, 2PC, and the operational cost of shardingsenior
- The seven acts: from CREATE TABLE to Citusjunior
- Acts 1–3 in depth: schema, indexes, and planner statisticsmiddle
- Acts 4–6 in depth: MVCC bloat, connection pooling, and safe migrationsmiddle
- Act 7 in depth: sharding, co-location, and the seven-tier tradeoff cascademiddle
- Observability, anti-patterns, and production triagesenior
- Raft roles, terms, and why majority quorums prevent split brainjunior
- How Raft replicates a log entry and decides it is safe to commitmiddle
- Raft leader election: timeouts, voting rules, and the four safety propertiesmiddle
- Raft in the real world: partitions, slow disks, and client routingmiddle
- Raft extensions: pre-vote, learners, snapshots, and linearizable readssenior
- Raft in production: membership changes, Multi-Raft, and observabilitysenior
- Where data fetching happens — and why it decides LCPjunior
- Fetch waterfalls — diagnosis and the Promise.all curemiddle
- React Server Components and Suspense streamingmiddle
- Client-side cache: TanStack Query, SWR, and stale-while-revalidatemiddle
- LCP, prefetch, and race conditions in interactive fetchingmiddle
- Senior internals: RSC payload, caching layers, and production failure modessenior
- The three-way handshakejunior
- Sequence numbers and connection statemiddle
- DNS: what it does and why it existsjunior
- The resolver walk: referrals, record types, and gluemiddle
- TTL, caching, and DNS propagationmiddle
- The 1-RTT handshake: key shares and ECDHEmiddle
- Session resumption and 0-RTTmiddle
- WebSocket: the HTTP upgrade handshakejunior
- WebSocket frame format: opcodes, masking, fragmentationmiddle
- WebSocket backpressure: when clients can''''t keep upmiddle
- Reconnection: jittered backoff, thundering herd, message resumptionsenior
- WebSocket at scale: HTTP/2 multiplexing, permessage-deflate, C10Msenior
- WebSocket in production: proxies, security, and distributed architecturesenior
- What reverse proxies dojunior
- Health checks, connection draining, and slow startmiddle
- Session affinity, consistent hashing, and the right fixmiddle
- Retry storms, circuit breakers, and load sheddingsenior
- Resilient LB architecture: anycast, zone-aware routing, and observabilitysenior
- Why QUIC and not TCP+TLSjunior
- Connection IDs and network migrationmiddle
- 0-RTT resumption and packet encryptionsenior
- DDoS: what it is and why it worksjunior
- Amplification attacks and state exhaustionmiddle
- Rate limiting: algorithms and architecturemiddle
- WAFs, firewalls, mTLS, and HSTSmiddle
- DNS cache poisoning and BGP hijackingsenior
- Defense-in-depth architecture and attack economicssenior
- DNS, TCP, TLS in sequence: where the milliseconds gomiddle
- Proxy intercepts and security gates: rate limiters, WAF, mTLSmiddle
- Alternate paths: QUIC 0-RTT, WebSocket upgrade, connection migrationmiddle
- Observability: distributed traces, USE/RED, and samplingsenior
- Resilience: cascading retries, circuit breakers, and error budgetssenior
- What the three signals are: logs, metrics, and tracesjunior
- Why structured logs exist: the diary vs the spreadsheetjunior
- The production log schema: fields every line must carrymiddle
- PII redaction and log injectionsenior
- OTel Logs Data Model and audit logs as a subsystemsenior
- SLI, SLO, and the error budget: reliability by the numbersjunior
- Error budget policy, latency SLOs, and composite journeysmiddle
- Production SLO failures, self-observability, security, and the big picturesenior
- The incident loop: from pager to postmortem to preventionmiddle
- Cache lines, struct layout, and false sharingmiddle
- SIMD, SoA vs AoS, and memory bandwidthmiddle
- Cache-oblivious algorithms, PGO, and production failuressenior
- GC in production: observability, security, edge cases, and fleet governancesenior
- Batching: amortize fixed cost per operationjunior
- The batching window: size and wait timemiddle
- Batching in Kafka and Postgresmiddle
- io_uring and observability of batchingmiddle
- From Nagle to io_uring: evolution of batchingmiddle
- Backpressure, failure isolation, and batch security in productionsenior
- CI enforcement and RUM: making budgets stickmiddle
- V8 JIT pipeline, HTTP priorities, and bundle securitysenior
- The performance loop: discipline, not a projectjunior
- Classify and fix: matching bottleneck families to remediesmiddle
- Observability stack and CI gates: catching regressions before they shipmiddle
- Incident to enforcement: SLO burn to verified fix in 35 minutesmiddle
- Culture, economics, and org-scale performancesenior
- At-most-once, at-least-once, exactly-once: the three delivery contractsjunior
- The three failure legs — where duplicates and losses actually happenmiddle
- Consumer-side dedup: the cheapest path to exactly-once processingmiddle
- Kafka exactly-once semantics: idempotent producer and transactionsmiddle
- SQS visibility timeout, DLQ, and the outbox patternmiddle
- Exactly-once in production: impossibility proof, hybrid patterns, and real incidentssenior
- What OAuth is and why passwords are not the answerjunior
- Authorization code flow with PKCEmiddle
- ID token validation and JWKS cache managementmiddle
- Refresh token rotation and scope-based least privilegemiddle
- Sender-constrained tokens: DPoP and mTLSsenior
- OAuth in production: audience attacks, observability, and real failuressenior
Something unclear?
Ask a question about this lesson. Questions are anonymous and go straight to the author to make the lesson better.