open atlas
↑ Back to track
JavaScript Engine internals JSE · 04 · 04

Speculation and guards

TurboFan bakes optimistic assumptions from feedback into fast code, protected by cheap guards — CheckMaps, CheckSmi, CheckBounds, CheckString. How guards gate the fast path, how inlining fuses small callees

JSE Senior ◷ 15 min
Level
FoundationsJuniorMiddleSenior

The fastest way to read point.x is a single machine instruction: load the value at a fixed byte offset from the object pointer. But that is only correct if point really has the shape TurboFan assumed when it compiled the code. So TurboFan does something audacious — it emits that one-instruction load, and in front of it puts a single cheap comparison: “is this object’s map still the one I expect?” If yes, the fast load runs. If no, the whole function bails. That comparison is a guard, and the entire art of speculative optimisation is making guards that almost never fail.

Optimism, protected

The previous lessons established the inputs: feedback says “this site always saw map M and Smi operands”, and TurboFan can specialise on that. But “always saw” is a statement about the past. JavaScript is free to hand the function something different on the next call. Speculative optimisation resolves this tension with a simple contract: assume the observed types, emit code that is only valid under that assumption, and protect every assumption with a guard that checks it cheaply before the fast code runs.

A guard is a tiny instruction sequence that tests the assumption and, on failure, transfers control out of the optimised code entirely (a deopt — lesson 05). The key property is asymmetry: when the guard passes — the overwhelmingly common case for stable code — it costs almost nothing (a compare and a not-taken branch), and the code behind it is maximally fast because it can assume the type with no further checks. The cost of correctness is paid only on the rare failing path.

The guards you will see in --trace-deopt reasons:

  • CheckMaps — “this object still has hidden class M”. Gates a property load/store at a fixed offset. Failure reason: wrong map.
  • CheckSmi / CheckNumber — “this value is a small integer” / “is a number”. Gates integer or float64 arithmetic with no tag checks or boxing. Failure reason: not a Smi, not a heap number.
  • CheckBounds — “this index is within the array’s length”. Gates a raw element load with no bounds re-check inside the loop. Failure reason: out of bounds.
  • CheckString, CheckInternalizedString, CheckHeapObject, and friends — similar guards for string and pointer assumptions.

Inlining: fusing guards and bodies

Ask yourself: if the optimiser can’t see past a call boundary, what happens to a guard the caller already proved, but the callee duplicates? Without inlining, you pay for it twice. A function call is a barrier: the optimiser can’t see across it, and the call itself costs a frame setup. Inlining removes the barrier for small, hot callees whose feedback names a known target. TurboFan splices the callee’s graph into the caller, and now the callee’s guards and the caller’s guards live in the same scope. The wins compound:

  • A CheckMaps the caller already proved is not repeated inside the inlined body (redundancy elimination across the inline boundary).
  • The callee’s pure computation can be scheduled together with the caller’s, hoisted out of shared loops, common-subexpressioned.
  • Escape analysis (next) can now see that an object the callee allocates and returns never actually escapes the combined function.

Polymorphic inlining handles up to ~4 known targets with an upfront map check that dispatches between the inlined bodies; beyond that the site is megamorphic and the call stays a generic, uninlined call.

Escape analysis: deleting the allocation

This is the optimisation that most surprises people, because it makes objects free. Escape analysis asks of each allocation: does any reference to this object escape the function — stored into a field that outlives the call, passed to a callee that might retain it, returned? If the answer is no — the object is purely a local scratchpad — TurboFan performs scalar replacement: it deletes the allocation entirely and replaces the object’s fields with ordinary SSA values that live in registers.

Consider a hot distance computation that builds a temporary point:

function dist(ax, ay, bx, by) {
  const d = { x: ax - bx, y: ay - by }; // temp object, never escapes
  return Math.sqrt(d.x * d.x + d.y * d.y);
}

Without escape analysis, every call allocates d on the heap, writes two fields, reads them back, and leaves garbage for the GC. With escape analysis, TurboFan proves d never escapes dist, deletes the allocation, and treats d.x and d.y as two register-resident values. The emitted code has zero heap allocation, no field stores, no field loads, and produces no GC pressure — the object existed only in the source text, not in the running machine code. This is enormous for the temporary points, iterators, options bags, and destructuring intermediates that pervade real code: written for clarity, compiled to nothing.

Guards and the optimisations they unlock
CheckMaps cost (pass)
1 compare + branch
Fast property load
1 mov at fixed offset
CheckSmi gates
untagged integer arithmetic
CheckBounds gates
raw element load, no re-check
Polymorphic inline cap
~4 targets
Escape analysis result
0 alloc for non-escaping obj
Scalar-replaced fields
live in registers
Failure path
deopt to interpreter
Quiz

Why is a CheckMaps guard cheap enough to put in front of every optimised property access?

Quiz

Which version of a temp object can escape analysis eliminate?

Order the steps

Order what optimised code does at a speculated, monomorphic property read.

  1. 1 Load the object's map pointer from its header
  2. 2 CheckMaps: compare it to the expected map M
  3. 3 Map matches: take the fast path
  4. 4 Load the property with one instruction at the precomputed offset
Why this works

Why guard rather than re-derive the type each time? Because re-deriving — a full type dispatch on every operation — is exactly the generic interpreter behaviour TurboFan exists to escape. The guard is a bet: pay one cheap check up front, then run a long stretch of code that assumes the answer and needs no further checks. For stable, monomorphic code the bet pays off on essentially every call, and the amortised cost of correctness rounds to zero. The bet only goes bad when the assumption breaks repeatedly — which is the deopt-loop of lesson 05.

Recall before you leave
  1. 01
    What is a guard, name the common ones, and why is the design asymmetric?
  2. 02
    How does inlining interact with guards and escape analysis?
  3. 03
    Explain escape analysis and scalar replacement with an example, and state its limit.
Recap

Speculative optimisation is how TurboFan turns past observations into fast code without giving up correctness. It assumes the types the feedback vector recorded, emits machine code valid only under those assumptions, and protects every assumption with a cheap guard: CheckMaps for object shape, CheckSmi/CheckNumber for numeric type, CheckBounds for array indices, plus string and pointer guards. The design is deliberately asymmetric — a passing guard is a single compare and a not-taken branch, so the fast path behind it (a one-instruction offset load, untagged arithmetic, a raw element read) is maximally fast, and real cost is paid only on the rare failing path, which deopts to the interpreter. Inlining fuses small, hot, known-target callees into the caller so their guards and bodies share a scope, enabling cross-boundary redundancy elimination, scheduling, and — crucially — escape analysis. Escape analysis identifies allocations that never escape the function and scalar-replaces them: the allocation is deleted and its fields become registers, so temporary points, iterators, and options bags compile to no allocation and no GC pressure at all. The one caveat: escape analysis only helps objects that provably stay local; anything returned, stored, or handed to an opaque callee escapes and is heap-allocated as usual. Now when you write a tight inner loop building a temporary {x, y} object for each iteration, you know whether that allocation is real or free — and what to check if GC pressure surprises you.

Practice

Start at the top. Tasks go easiest → hardest: recall a fact, apply it to a case, then a senior-level stretch. Open one, attempt it, then reveal.

recallapplystretch0 of 8 done
Connected lessons
appears again in184

Something unclear?

Ask a question about this lesson. Questions are anonymous and go straight to the author to make the lesson better.

shortcuts expand
search
K
prev piece
k
next piece
j
cycle tier
t
this menu
?
sources3
expand
  1. 01
  2. 02
  3. 03

Trademarks belong to their respective owners. Editorial reference only.