open atlas
↑ Back to track
JavaScript Engine internals JSE · 09 · 01

Capstone: optimize a hot path

One slow request handler, every layer of the engine in play. Profile it, read the deopt and IC traces, fix the shape, the boxing, the GC churn, and the microtask stall — then prove the win. The whole track, applied to one function.

JSE Senior ◷ 16 min
Level
FoundationsJuniorMiddleSenior

A scoring endpoint that used to clear 40k requests/second now barely manages 6k, and the only “change” was a refactor that “just tidied up the object construction.” Nothing looks wrong in the code. Everything you learned in this track is about to converge on one 30-line function — because the engine is doing five different slow things at once, and you can name and fix every one.

The method, before the function

Optimisation without measurement is superstition. The loop is always the same: profile to find the dominant cost, name the mechanism, make one change, re-measure. This track gave you the vocabulary for the “name the mechanism” step — which is where most engineers get stuck, guessing instead of reading the trace. By the end of this capstone you will have named and fixed every slow thing in this function — and know exactly which tool to reach for next time.

Here is the regressed handler. Read it as the engine would.

function scoreBatch(rows) {
  const out = [];
  for (const row of rows) {
    const e = {};
    e.id = row.id;
    if (row.name) e.name = row.name;       // conditional field
    e.score = row.weight * 1e9 + row.bonus; // overflows Smi range
    if (row.flagged) e.flag = true;          // late conditional field
    out.push(e);
    log(`scored ${e.id}`);                    // synchronous console in the hot loop
  }
  return out;
}

Layer 1 — the shape (units 02–03)

--trace-ic shows the access site that reads .score downstream going megamorphic. The cause is the conditional fields: name and flag are added only sometimes, so objects walk different transition-tree branches and land on different hidden classes (unit 03). The IC at the consumer sees five-plus maps and falls back to the generic lookup — tens of cycles instead of one. When you see megamorphic or GENERIC in a trace, your first question should be: which call site produces objects with inconsistent shapes?

Fix: initialise every property unconditionally, in a fixed order, so all objects share one Map.

const e = {
  id: row.id,
  name: row.name ?? null,
  score: 0,                 // set below
  flag: row.flagged === true,
};

Layer 2 — the boxing and the deopt (units 02, 04)

--trace-deopt shows scoreBatch deoptimising on a CheckSmi guard. row.weight * 1e9 overflows the Smi range (±2³¹), so the result becomes a HeapNumber — a heap-allocated box (unit 02). TurboFan had speculated Smi from early feedback; the first overflow fails the guard and triggers a deopt, and because it recurs every batch, you get a deopt loop (unit 04): optimise → deopt → re-optimise, never staying fast.

Fix: stop pretending the value is an integer. Compute in double from the start so the optimiser specialises on Float64 and never guards on Smi — or, if the consumer is numeric-heavy, accumulate scores into a Float64Array instead of object fields, eliminating the box entirely.

Cost of each mechanism on this path
Megamorphic property load vs monomorphic
~10–50×
Deopt loop (optimise/deopt churn)
stays interpreted
HeapNumber box vs Smi
alloc + deref vs 1 instr
Minor GC triggered by per-row allocation
sub-ms, but frequent
Synchronous log() in a 100k loop
blocks the turn

Layer 3 — the GC churn (unit 06)

Even with shapes fixed, allocating one out object per row floods new space; the bump-pointer allocator fills a semi-space and triggers a Scavenge (unit 06) far more often than necessary. Minor GC is cheap individually but death by a thousand cuts at this volume, and survivors get promoted to old space, eventually forcing a major GC.

Fix: reduce allocation. If out feeds a reducer, fold rows directly instead of materialising an array of objects; if the array is required, pre-size it (new Array(rows.length) is fine here because you fill every slot, keeping it PACKED). Fewer, longer-lived allocations beat a storm of short-lived ones.

Layer 4 — the scheduling (unit 07)

The log() call is synchronous I/O inside the loop. Beyond its own cost, in a request context it interleaves with the event loop: a flood of synchronous work on the stack delays the microtask checkpoint and starves other handlers (unit 07). The fix is not “make logging async” — it is don’t log per row on a hot path. Aggregate and emit once, or sample.

Layer 5 — prove it, then guard it (unit 08)

Measure the rewrite against the original on warmed-up code, consuming the result so dead-code elimination can’t delete your benchmark (unit 08), with allocation kept out of the timed region and the engine recorded. Then add a CI microbenchmark asserting a throughput floor, plus a unit test that counts distinct Object.keys signatures from scoreBatch and fails if it ever exceeds one. The shape regression that started this incident now cannot ship silently again.

Quiz

Why fix the hidden-class instability before chasing the deopt or the GC churn?

Quiz

`row.weight * 1e9 + row.bonus` triggered a deopt loop. What is the precise mechanism?

Order the steps

Order the optimisation method for this incident.

  1. 1 Profile to find the dominant cost (--prof, --trace-deopt, --trace-ic, heap snapshot)
  2. 2 Name the engine mechanism behind it (shape / boxing / GC / scheduling)
  3. 3 Make the smallest change that targets that mechanism
  4. 4 Re-measure on warmed-up code with the result consumed
  5. 5 Add a regression gate so the win cannot silently revert
More practice

Apply this to your own code: pick one function your profiler flags as hot. Before touching it, write down which of the five layers you expect to be the cost, then run --trace-deopt and a heap snapshot to check. Most engineers are wrong about which layer dominates — that gap between intuition and the trace is exactly what this track was built to close.

Recall before you leave
  1. 01
    The handler is slow. Walk the full diagnostic method end to end.
  2. 02
    Why does conditional property initialisation regress a hot consumer, and what's the fix?
  3. 03
    When is rewriting the hot path in WASM/Rust actually justified?
Recap

This capstone put the whole track on one function. A regressed scoring handler was slow for five overlapping reasons, each nameable from this track: conditional property initialisation split objects across hidden classes and drove a consumer IC megamorphic (units 02–03); an integer that overflowed the Smi range became a HeapNumber and triggered a TurboFan deopt loop (units 02, 04); per-row object allocation churned new space into frequent Scavenges (unit 06); a synchronous log per iteration stalled the microtask checkpoint (unit 07). The method is invariant: profile to find the dominant cost, name the exact mechanism, change one thing, re-measure on warmed-up code with the result consumed, and lock the win behind a regression gate (unit 08). Fix shape first, because a stable shape is the precondition for the optimiser to produce trustworthy code and trustworthy measurements. Now when you see a slow handler, you will reach for the trace first, name the mechanism precisely, and fix one layer at a time — because that is what turns guessing into engineering.

Practice

Start at the top. Tasks go easiest → hardest: recall a fact, apply it to a case, then a senior-level stretch. Open one, attempt it, then reveal.

recallapplystretch0 of 8 done
Connected lessons
appears again in184

Something unclear?

Ask a question about this lesson. Questions are anonymous and go straight to the author to make the lesson better.

Apply this

Put this lesson to work on a real build.

shortcuts expand
search
K
prev piece
k
next piece
j
cycle tier
t
this menu
?
sources3
expand
  1. 01
  2. 02
  3. 03

Trademarks belong to their respective owners. Editorial reference only.