U+055F ARMENIAN ABBREVIATION MARK has Word_Break=MidLetter and is
Case_Ignorable in the UCD, but it was missing from the hardcoded list,
so the Greek final sigma context scan stopped at it instead of skipping
it. This affects String.downcase/2 and String.capitalize/2 in :greek
mode on both sides of the Final_Sigma rule.
Assisted-by: Claude Code:claude-opus-5
Signed-off-by: jaideeppyne <jaideeppyne1997@gmail.com>
trim_trailing/1 walked the string with binary_part/3, which allocated the
three byte lookahead plus a fresh prefix on every step even though every
prefix but the last one is discarded, and paid one allocation even when
there was nothing to trim.
Carry the position instead and slice once at the end. The six one-byte
whitespace codepoints, which are the common case, now need no lookahead at
all, so the lookahead tables only hold the multi-byte ones.
Dispatching on the byte with :binary.at/2 instead of matching the string is
what makes the common path free. Matching turns the argument into a match
context, and a clause returning the string unchanged then has to materialize
it again, which is the allocation we are trying to avoid.
Measured on OTP 29, ns/op and words allocated per call:
before after
100B, nothing to trim 47.6 / 8w 16.0 / 0w
1KB, nothing to trim 47.4 / 8w 16.2 / 0w
100B + one space 73.2 / 21w 41.4 / 5w
100B + four spaces 104.4 / 34w 83.2 / 5w
100B + NBSP 71.8 / 21w 60.9 / 13w
100B + 64 spaces 754.8 / 294w 887.8 / 5w
Trimming every line of lib/elixir/lib/kernel.ex drops from 0.442ms to
0.134ms, and 2000 lines ending in a newline from 0.282ms to 0.094ms.
The trade-off is the last row. The old code advanced up to three bytes per
iteration through the lookahead table where this one advances one byte at a
time, so runs longer than about five whitespace bytes lose time, 15% to 20%
on a 64 byte run. Allocation on that run goes from O(n) to O(1), and long
trailing runs are rare next to lines ending in a single newline or in
nothing at all.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ensure identifiers that require normalization are checked again
after normalization so decomposed source cannot normalize into
unsupported codepoints.
Closes#15419.
* update support for uts39 from unicode 15
* follow uts39's recco that it's not necessary to require
idents to be single-script (they call out proglang idents,
reference the new uts55-5). We use a heuristic derived from
the concept of identifier chunks from uts55-5, to allow
idents like foo_bar_baz where each chunk around the _ can be
single-or-highly-restrictive
* provide directional confusability detection, by reversing
spans of direction-changed chars in idents for bidi_skeleton,
see issue #12929
When capitalize/2 was written, Erlang did not provide titlecase
functions. capitalize/2 also downcases the rest of the string
while Erlang doesn't, which is a common contention point.
Given our titlecase implementation requires 40kb of additional
code in .beam files, it makes sense to unify it with Erlang's
(which we have already done for most non-critical functionality
in unicode).
* nfc by default
* suggest nfkc on unexpected tokens when nfkc would have worked
* support additional normalizations (just micro/mu for now)
* parser, formatter tests
* tweak Macro.inner_classify(atom) not to rely on not_nfc causing error
The patch computes byte lookups based on the prefix. For example,
Á, É, etc all have the same prefix <<195>>, so they are lumped
together for lookup and then we just do a byte lookup later. We
tried doing the byte lookup on 64-element tuple (since the byte
is always within 0b10000000 and 0b10111111) but that's slower,
especially because we need to check the byte range for invalid
Unicode, so instead the last byte lookup is a case. Grouping the
top-level lookup makes the cost of a miss 3x cheaper albeit a
hit is 10% more expensive and reduces bytecode size.
Erlang ships with its own embedding of the Unicode
Codebase for a couple releases and this release changes
Elixir to depend on it in order to compute grapheme
clusters.
The Erlang implementation was up to 2x faster in low
codepoints (such as latin1) while the Elixir one could
be faster up to 3x in high codepoints (such as emoji)
so at the end the performance results are roughly the
same. As a benefit, we no longer need to ship our copy
of the grapheme cluster algorithm, which would take up
to 250kB in disk and more than 15 seconds to compile.
Note we still keep our own String downcase and upcase
algorithms, as our version is considerably more efficient
on all cases since it works exclusively with binaries
(more than 5x faster).