Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 37 additions & 0 deletions changelog.d/11643-inline-number-to-string.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
perf(codegen): number-to-string builds a small integer's text at the call site, and its result keeps its string proof (#10762).

- `String(n)`, `` `${n}` ``, `n.toString()` and `"" + n` on an operand the type
analysis proves numeric now compute the SSO bits of an integer in
`-9999..=99999` inline, with no call. The digits come from a fixed-point
split (`abs * ceil(2^32 / 10^4)`, then `* 10` per digit, exact for every
`abs < 100000`), about 30 branch-free instructions. Every other value
(fractions, `NaN`, the infinities, larger integers, and a declared `number`
that holds something else at run time) takes the original runtime call on
a cold arm, so the text is unchanged, and the bits match
`small_integer_sso_bits` exactly. An operand without a numeric proof keeps
the plain call, so the code is emitted only where it is expected to fire.
- `const s = String(n)`, `` `${n}` `` and `"" + n` (a `+` with a string-literal
operand) record a runtime-derived `String` proof for `s`. Before, the local
lost the proof its initializer carries, and `s.charCodeAt(i)`, `s.length`
and the other string lowerings fell to the generic method site, although
`String(n).charCodeAt(i)` written inline took the fast path. A declared
`string` is still not a proof (#7837).
- The inline `charCodeAt` reads an all-ASCII SSO receiver's byte straight out
of the value. Before, an SSO receiver, which is what every short number's text
now is, went to the slow arm, which materialized a heap copy through the
intern table (about 175 instructions) to read one byte.

Measured (`callgrind`, instructions per iteration fitted over 20k→120k, x86-64,
`PERRY_NO_AUTO_OPTIMIZE=1`, `k = i % 1000`, the text consumed by
`charCodeAt(0) + length`): `String(n)` 449 → 169, `` `${n}` `` 462 → 169,
`String(-k)` 644 → 199, a fraction 1,251 → 987, and a 7-digit integer (a heap
string) 623 → 560 over 1M→3M. Consumed only by `.length`: `String(n)`
163 → 131, `"" + n` 186 → 133, `` `${n}` `` 172 → 131, `n.toString()`
160 → 133. A bare loop is 23.

Tests: `number_to_string_inline_tests` (each spelling emits the inline arm; an
`any` operand does not; `s.charCodeAt` on each coerced local reaches the SSO
arm), sabotage-checked, and `test_gap_10762_inline_number_to_string` (every
integer in the inline range against a reference, the range edges, lying
annotations, SSO `charCodeAt` including non-ASCII and bad indexes;
byte-identical to Node). No version bump.
44 changes: 44 additions & 0 deletions changelog.d/11643-map-own-override-flag-builtin-installs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
perf(runtime): populating `globalThis` no longer puts every `Map`/`Set`/`Date` builtin call on the slow side of the own-override guard (#10697).

The #10943 guard in front of a proven `Map`/`Set`/`Date` builtin call answers
"no own override" from one load of `PERRY_OWN_NAMED_PROP_INSTALLED`, and asks
the authoritative `hasOwn` predicate only once any non-ordinary cell has taken
a named property. The runtime's own lazy `globalThis` population set that flag:
it installs statics on constructor intrinsics such as `%TypedArray%` (closure
cells), `constructor` on `Array.prototype` (an array cell) and aliases such as
`Number.parseFloat`, and each of those stores passes the exotic-store gauntlet
that arms it. Any program that referenced `globalThis` or a lazily installed
global then paid about 600 instructions of `js_receiver_may_own_named_method`
→ `js_object_has_own` (with a key-string coercion) on every `m.get(k)` and
`m.set(k, v)`. That is the "count by category" shape #10697 measured at 4.2x
node.

- The runtime's own builtin definitions (`define_builtin_data_property` and all
of `populate_global_this_builtins`) run inside `as_builtin_definition`. Inside
it, an install arms the exported flag only when its owner is a Map, Set or
Date cell, or its header cannot be read. Those are the only receivers whose
guard answer comes from the flag: an array answers from its own header and
named-property storage, and a declared-`Map` receiver of any other kind is
brand-checked into generic dispatch. User installs arm exactly as before.
- The universal dispatcher's `own_user_method_value` now reads a separate
internal flag that every install still arms, builtin or not, so its answers
are unchanged.
- Small-map lookups (≤ 8 entries) answer a bit-identical key of any type from
the inlined hot lane. A Map never holds two SameValueZero-equal keys and
stored keys are normalized, so an identity hit is the match. Before, only a
plain-number key could use that lane, and every string lookup paid the
out-of-line call to `find_key_index_cold` first.

Measured (`callgrind`, instructions per iteration fitted over 20k→120k, x86-64,
`PERRY_NO_AUTO_OPTIMIZE=1`, program references `globalThis`): four constant
string keys, get-or-default then set, 1,866 → 373; a plain `m.get(CATS[i & 3])`
947 → 206, of which 111 is the array read itself. Output identical to Node.

Tests: `a_builtin_install_on_an_intrinsic_does_not_arm_the_guard` (272 arms
from one intrinsic install before; 0 now, while installs onto a Map, builtin
or user, still arm), `a_small_map_answers_an_identical_key_without_the_cold_path`
(8 of 8 identical lookups went cold before; 0 now, and content matches still
go cold), both sabotage-checked, and
`test_gap_10697_own_override_after_global_population` (own overrides on
Map/Set/Date/Array still win after population, and the populated builtins still
work). No version bump.
20 changes: 20 additions & 0 deletions changelog.d/11643-random-uuid-format.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
perf(uuid): `randomUUID()` formats without bounds checks or a UTF-8 re-validation (#10523).

#10523's main cost, one `getrandom` system call per UUID, was fixed by the
per-thread entropy cache (#11351): 120,000 UUIDs now make 944 calls, one per 128,
the batching Node uses. This removes most of what remained in user space:

- `Hyphenated::new` wrote each byte's two hex digits at a running index with a
per-byte hyphen test, so every store was bounds-checked. A constant table of
digit positions unrolls to straight-line stores.
- `js_crypto_random_uuid` and `js_crypto_random_uuidv7` copy the 36 bytes
through the new `Hyphenated::as_bytes`. `as_str` validated known-ASCII bytes
as UTF-8 on every UUID only for the string constructor to copy them.

Measured (`callgrind`, whole-process instructions ÷ UUIDs, `node:crypto`
`randomUUID()`): 1,085 → 1,005 per UUID at 200k, 1,119 → 933 at 1M. The two
hot functions fell from about 367 to 290 instructions per UUID, and the 56 of
UTF-8 validation are gone.

Tests: `hyphenated_layout_is_exact` pins the digit order and hyphen positions
byte for byte, and that `as_bytes` equals `as_str`'s bytes. No version bump.
27 changes: 27 additions & 0 deletions crates/perry-codegen/src/expr/logical_collections.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1188,13 +1188,40 @@ pub(crate) fn lower(ctx: &mut FnCtx<'_>, expr: &Expr) -> Result<String> {
// and `${i}` allocate nothing for it. The result is therefore
// SSO-or-heap, not heap — see `proven_heap_string_operand`.
Expr::StringCoerce(operand) => {
let numeric = crate::type_analysis::is_numeric_expr(ctx, operand);
let v = lower_expr(ctx, operand)?;
// A number operand's small-integer text is built inline (#10762).
if numeric {
return crate::expr::number_to_string_inline::emit_number_to_string_inline(
ctx,
&v,
|ctx| {
Ok(ctx
.block()
.call(DOUBLE, "js_string_coerce_box", &[(DOUBLE, &v)]))
},
);
}
Ok(ctx
.block()
.call(DOUBLE, "js_string_coerce_box", &[(DOUBLE, &v)]))
}
Expr::TemplateStringCoerce(operand) => {
let numeric = crate::type_analysis::is_numeric_expr(ctx, operand);
let v = lower_expr(ctx, operand)?;
if numeric {
return crate::expr::number_to_string_inline::emit_number_to_string_inline(
ctx,
&v,
|ctx| {
Ok(ctx.block().call(
DOUBLE,
"js_template_string_coerce_box",
&[(DOUBLE, &v)],
))
},
);
}
// S2: a string operand is answered inline; see `ic_fast_split.rs`.
Ok(crate::expr::ic_fast_split::emit_template_string_coerce(
ctx, &v,
Expand Down
3 changes: 3 additions & 0 deletions crates/perry-codegen/src/expr/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3116,6 +3116,9 @@ mod math_simple;
pub(crate) mod method_site;
mod misc_methods;
mod new_dynamic;
pub(crate) mod number_to_string_inline;
#[cfg(test)]
mod number_to_string_inline_tests;
mod objects_arrays_lit;
pub(crate) mod os_uri_dates;
pub(crate) mod property_get;
Expand Down
125 changes: 125 additions & 0 deletions crates/perry-codegen/src/expr/number_to_string_inline.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
//! Inline small-integer number-to-string (#10762).
//!
//! `String(n)`, `` `${n}` ``, `n.toString()` and `"" + n` on a number all
//! reach `perry-runtime`'s `small_integer_sso_bits` for an integer in
//! `-9999..=99999`: the answer is an SSO immediate, computed with no
//! allocation. What was left was the call itself and the tag ladder in front
//! of it. This emits the same computation at the call site for operands the
//! type analysis says are numbers, and keeps the runtime call on a cold arm
//! for everything else, including a declared-`number` slot that holds
//! something else at run time.
//!
//! The bits are identical to the runtime's, so the two arms are
//! interchangeable: most significant digit in byte 0, a leading `-` below the
//! digits, and the length at `SHORT_STRING_LEN_SHIFT`. `-0` converts to `0`
//! and compares equal to `0.0`, so it prints `"0"` like the runtime.

use super::FnCtx;
use crate::nanbox::{double_literal, i64_literal, SHORT_STRING_TAG};
use crate::types::{DOUBLE, I1, I32, I64};

/// Bit offset of an SSO string's length byte. Must match
/// `perry-runtime::value::tags::SHORT_STRING_LEN_SHIFT`.
const SHORT_STRING_LEN_SHIFT: u64 = 40;

/// Emit the SSO bits of `v` inline when it is an integral double in
/// `-9999..=99999`, and `slow(ctx)` otherwise. Returns the merged NaN-boxed
/// string. `slow` runs with the builder positioned in the cold block and must
/// return a DOUBLE.
pub(crate) fn emit_number_to_string_inline(
ctx: &mut FnCtx<'_>,
v: &str,
slow: impl FnOnce(&mut FnCtx<'_>) -> anyhow::Result<String>,
) -> anyhow::Result<String> {
let int_idx = ctx.new_block("num2str.int");
let fast_idx = ctx.new_block("num2str.fast");
let slow_idx = ctx.new_block("num2str.slow");
let merge_idx = ctx.new_block("num2str.merge");
let int_label = ctx.block_label(int_idx);
let fast_label = ctx.block_label(fast_idx);
let slow_label = ctx.block_label(slow_idx);
let merge_label = ctx.block_label(merge_idx);

// Range first: `fptosi` of an out-of-range value is poison, so it runs
// only behind this branch. Every comparison with a NaN — a NaN-boxed
// non-number included — is false.
{
let blk = ctx.block();
let lo = blk.fcmp("oge", v, &double_literal(-9_999.0));
let hi = blk.fcmp("ole", v, &double_literal(99_999.0));
let in_range = blk.and(I1, &lo, &hi);
blk.cond_br(&in_range, &int_label, &slow_label);
}

ctx.current_block = int_idx;
let int = {
let blk = ctx.block();
let int = blk.fptosi(DOUBLE, v, I32);
let back = blk.sitofp(I32, &int, DOUBLE);
let integral = blk.fcmp("oeq", &back, v);
blk.cond_br(&integral, &fast_label, &slow_label);
int
};

ctx.current_block = fast_idx;
let fast = {
let blk = ctx.block();
let neg = blk.icmp_slt(I32, &int, "0");
let negated = blk.sub(I32, "0", &int);
let abs32 = blk.select(I1, &neg, I32, &negated, &int);
let abs = blk.zext(I32, &abs32, I64);
// Five digits, most significant first, by fixed-point division:
// `abs * ceil(2^32 / 10^4)` puts `abs / 10^4` in the high word and
// the scaled remainder in the low word, and each `* 10` of the low
// word shifts the next digit up. Exact for every `abs < 100000`
// (checked exhaustively), one multiply per digit where the plain
// `/ 10`, `% 10` pairs cost about five instructions each. Leading
// zeros are digits too, so they are shifted out below.
let low_mask = i64_literal(0xFFFF_FFFF);
let mut scaled = blk.mul(I64, &abs, "429497");
let mut packed = blk.lshr(I64, &scaled, "32");
for byte in 1..5u32 {
let low = blk.and(I64, &scaled, &low_mask);
scaled = blk.mul(I64, &low, "10");
let digit = blk.lshr(I64, &scaled, "32");
let shifted = blk.shl(I64, &digit, &(byte * 8).to_string());
packed = blk.or(I64, &packed, &shifted);
}
let ascii = blk.or(I64, &packed, &i64_literal(0x30_3030_3030));
let mut ndig = "1".to_string();
for bound in ["10", "100", "1000", "10000"] {
let ge = blk.icmp_uge(I64, &abs, bound);
let one = blk.zext(I1, &ge, I64);
ndig = blk.add(I64, &ndig, &one);
}
let zeros = blk.sub(I64, "5", &ndig);
let shift = blk.shl(I64, &zeros, "3");
let payload = blk.lshr(I64, &ascii, &shift);
let with_sign = blk.shl(I64, &payload, "8");
let with_sign = blk.or(I64, &with_sign, &(b'-' as u64).to_string());
let signed_len = blk.add(I64, &ndig, "1");
let payload = blk.select(I1, &neg, I64, &with_sign, &payload);
let len = blk.select(I1, &neg, I64, &signed_len, &ndig);
let len = blk.shl(I64, &len, &SHORT_STRING_LEN_SHIFT.to_string());
let bits = blk.or(I64, &payload, &len);
let bits = blk.or(I64, &bits, &i64_literal(SHORT_STRING_TAG));
let boxed = blk.bitcast_i64_to_double(&bits);
blk.br(&merge_label);
boxed
};

ctx.current_block = slow_idx;
let slow_val = slow(ctx)?;
let slow_end = ctx.block().label.clone();
ctx.block().br(&merge_label);

ctx.current_block = merge_idx;
let fast_label_end = ctx.block_label(fast_idx);
Ok(ctx.block().phi(
DOUBLE,
&[
(fast.as_str(), fast_label_end.as_str()),
(slow_val.as_str(), slow_end.as_str()),
],
))
}
Loading
Loading