Commit Graph
134 Commits
Author SHA1 Message Date
Jaideep Pyne a164205866 Add U+055F to the case ignorable set (#15812)
U+055F ARMENIAN ABBREVIATION MARK has Word_Break=MidLetter and is
Case_Ignorable in the UCD, but it was missing from the hardcoded list,
so the Greek final sigma context scan stopped at it instead of skipping
it. This affects String.downcase/2 and String.capitalize/2 in :greek
mode on both sides of the Final_Sigma rule.

Assisted-by: Claude Code:claude-opus-5
Signed-off-by: jaideeppyne <jaideeppyne1997@gmail.com>
2026-08-30 19:02:30 +02:00
Alexander Gubarev 093d4c8873 Fix final sigma handling in :greek (#15791) 2026-08-24 16:05:49 +02:00
Daniel KukulaandClaude Opus 5 11b7f08b0b Trim trailing whitespace without slicing per step (#15780)
trim_trailing/1 walked the string with binary_part/3, which allocated the
three byte lookahead plus a fresh prefix on every step even though every
prefix but the last one is discarded, and paid one allocation even when
there was nothing to trim.

Carry the position instead and slice once at the end. The six one-byte
whitespace codepoints, which are the common case, now need no lookahead at
all, so the lookahead tables only hold the multi-byte ones.

Dispatching on the byte with :binary.at/2 instead of matching the string is
what makes the common path free. Matching turns the argument into a match
context, and a clause returning the string unchanged then has to materialize
it again, which is the allocation we are trying to avoid.

Measured on OTP 29, ns/op and words allocated per call:

                          before        after
  100B, nothing to trim   47.6 / 8w    16.0 / 0w
  1KB, nothing to trim    47.4 / 8w    16.2 / 0w
  100B + one space        73.2 / 21w   41.4 / 5w
  100B + four spaces     104.4 / 34w   83.2 / 5w
  100B + NBSP             71.8 / 21w   60.9 / 13w
  100B + 64 spaces       754.8 / 294w 887.8 / 5w

Trimming every line of lib/elixir/lib/kernel.ex drops from 0.442ms to
0.134ms, and 2000 lines ending in a newline from 0.282ms to 0.094ms.

The trade-off is the last row. The old code advanced up to three bytes per
iteration through the lookahead table where this one advances one byte at a
time, so runs longer than about five whitespace bytes lose time, 15% to 20%
on a 64 byte run. Allocation on that run goes from O(n) to O(1), and long
trailing runs are rare next to lines ending in a single newline or in
nothing at all.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 00:57:08 +02:00
Thomas Cioppettini 30af9521cf Remove duplicate pattern compilation in Unicode data processing (#15776)
Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Thomas Cioppettini <544875+tomciopp@users.noreply.github.com>
2026-08-21 09:14:19 +02:00
José Valim ec50a88ee4 Validate normalized tokenizer identifiers
Ensure identifiers that require normalization are checked again
after normalization so decomposed source cannot normalize into
unsupported codepoints.

Closes #15419.
2026-06-18 10:55:09 +02:00
José Valim 175f54869b Update Unicode to version 17.0.0 (#14760)
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance

Please read the release notes by visiting
<http://www.unicode.org/versions/Unicode17.0.0/>.
2025-09-10 09:09:44 +02:00
Jonatan Männchen 980500551c Inline Licence Info & Copyright for non-test files (#14256) 2025-02-07 21:11:09 +01:00
José Valim d3cef1f32c Type checking of protocols in for-comprehensions (#14124) 2024-12-29 11:49:04 +01:00
maintenance-beam-app 13d3cc0a3f Update Unicode to version 16.0.0 (#14024)
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance

Before merging, please read the release notes by visiting
<http://www.unicode.org/versions/Unicode16.0.0/>
and assess if additional changes are necessary in the code base.
2024-11-29 10:01:27 +01:00
José Valim c08a9b371c Update CHANGELOG 2024-07-02 15:21:25 +02:00
José Valim 692b13dcfe Remove unused import 2024-07-02 14:56:21 +02:00
José Valim f01dc17d85 Optimize reversing logic
Rule #1:

    Enum.reverse([1, 2, 3]) ++ [4, 5, 6]

is equivalent but slower than

    Enum.reverse([1, 2, 3], [4, 5, 6])

Rule #2:

    Enum.reverse([1, 2, 3] ++ [4, 5, 6])

is equivalent but slower than:

    Enum.reverse([4, 5, 6], Enum.reverse([1, 2, 3]))

Rule #3

    Enum.reverse(Enum.reverse([1, 2, 3]))

is the same as:

    [1, 2, 3]
2024-07-02 14:51:49 +02:00
José Valim 9924afff5d Remove highly restrictive scriptset support 2024-07-02 14:41:57 +02:00
Luc Fueston c83334e5f2 Directional confusability protection & allowing script mixing in identifiers when separated by underscores (#13693)
* update support for uts39 from unicode 15

* follow uts39's recco that it's not necessary to require
  idents to be single-script (they call out proglang idents,
  reference the new uts55-5). We use a heuristic derived from
  the concept of identifier chunks from uts55-5, to allow
  idents like foo_bar_baz where each chunk around the _ can be
  single-or-highly-restrictive

* provide directional confusability detection, by reversing
  spans of direction-changed chars in idents for bidi_skeleton,
  see issue #12929
2024-07-02 14:07:20 +02:00
José Valim 47079bc0b2 Add more deprecations scheduled to v1.18 2024-05-26 21:44:10 +02:00
EksperimentalandMaintenance App a6b68dd684 Update Unicode to version 15.1.0 (#12947)
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance

Before merging, please read the release notes by visiting
<http://www.unicode.org/versions/Unicode15.1.0/>
and assess if additional changes are necessary in the code base.

Co-authored-by: Maintenance App <maintenance-beam@autistici.org>
2023-09-20 09:16:13 +02:00
José Valim 76a92b5a80 Deprecate String.capitalize/2 in favor of :string.titlecase/1
When capitalize/2 was written, Erlang did not provide titlecase
functions. capitalize/2 also downcases the rest of the string
while Erlang doesn't, which is a common contention point.

Given our titlecase implementation requires 40kb of additional
code in .beam files, it makes sense to unify it with Erlang's
(which we have already done for most non-critical functionality
in unicode).
2023-08-09 16:15:41 +02:00
Andrea Leopardi dd295f4bc7 Fix spelling of "hard-coded" 2023-07-09 16:08:12 +02:00
José Valim 7454333fd0 Require pin variable when accessing variable inside binary size in match, closes #12588 2023-05-26 16:35:36 +02:00
maintenance-beam-app e8d8e7d693 Update Unicode to version 15.0.0 (#12143)
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance

Before merging, please read the release notes by visiting
<http://www.unicode.org/versions/Unicode15.0.0/>
and assess if additional changes are necessary in the code base.
2022-09-18 00:04:19 +02:00
sabiwara 29bf449f3f Evaluate formating charlists as ~c sigils (#12064) 2022-08-10 06:47:00 +02:00
José Valim fbf94613f3 Fix regression on identifier\\ being interpreted as a single identifier 2022-06-30 15:52:18 +02:00
José Valim 850c7ddfe4 Normalize unicode characters once in tokenizer 2022-05-23 12:14:39 +03:00
José Valim 25212f2db1 Do not attempt to suggest starting tokens
It is hard to suggest because it is not possible to
know if the user intended an operator, alias, or
identifier, so we can give incorrect feedback.
2022-05-23 12:14:39 +03:00
Luc Fueston 31cbdbd6e0 nfc and additional normalizations for identifiers (#11859)
* nfc by default
* suggest nfkc on unexpected tokens when nfkc would have worked
* support additional normalizations (just micro/mu for now)
* parser, formatter tests
* tweak Macro.inner_classify(atom) not to rely on not_nfc causing error
2022-05-23 09:44:44 +02:00
John Bampton 83467dec72 Remove trailing whitespace (#11782) 2022-04-25 15:25:31 +02:00
José Valim f067abaf11 Remove duplication in unicode tokenizer 2022-04-09 09:46:56 +02:00
José Valim bd5ee19108 Simplify process of copying SpecialCasing.txt 2022-02-13 09:01:05 +01:00
Eksperimental d806a6858f Improve Unicode module update instructions (#11622) 2022-02-13 08:54:24 +01:00
José Valim 2676f8b221 Allow only single- and highly restricted mixed-scripts in identifiers (UTS39 C3) (#11621) 2022-02-11 20:44:06 +01:00
José Valim bee0bb56be Do not emit warnings when formatting code
Closes #11595.
2022-01-26 13:37:53 +01:00
José Valim 851ac27272 Run unicode security linting on more tokens 2022-01-21 10:50:02 +01:00
Luc Fueston cba50b7a28 Warn on confusable non-ascii identifiers (UTS 39, C2) (#11582) 2022-01-21 09:37:25 +01:00
Luc Fueston 1b80c8c838 Don't allow restricted characters in identifiers (UTS 39, C1) (#11580) 2022-01-17 19:42:36 +01:00
José Valim a8b72c759a Group Unicode upcase/downcase by prefix (#11310)
The patch computes byte lookups based on the prefix. For example,
Á, É, etc all have the same prefix <<195>>, so they are lumped
together for lookup and then we just do a byte lookup later. We
tried doing the byte lookup on 64-element tuple (since the byte
is always within 0b10000000 and 0b10111111) but that's slower,
especially because we need to check the byte range for invalid
Unicode, so instead the last byte lookup is a case. Grouping the
top-level lookup makes the cost of a miss 3x cheaper albeit a
hit is 10% more expensive and reduces bytecode size.
2021-10-14 14:44:13 +02:00
Eksperimental 9b80f8a40b Upgrade Unicode database to v14.0.0 (#11308)
Closes #11288.
2021-10-12 18:05:10 +02:00
José Valim b2491fd839 Depend on Erlang for computing grapheme clusters (#11024)
Erlang ships with its own embedding of the Unicode
Codebase for a couple releases and this release changes
Elixir to depend on it in order to compute grapheme
clusters.

The Erlang implementation was up to 2x faster in low
codepoints (such as latin1) while the Elixir one could
be faster up to 3x in high codepoints (such as emoji)
so at the end the performance results are roughly the
same. As a benefit, we no longer need to ship our copy
of the grapheme cluster algorithm, which would take up
to 250kB in disk and more than 15 seconds to compile.

Note we still keep our own String downcase and upcase
algorithms, as our version is considerably more efficient
on all cases since it works exclusively with binaries
(more than 5x faster).
2021-06-01 12:44:55 +02:00
José Valim 5872d4947e Reduce size of Unicode module in 33% by not duplicating codepoints 2021-05-25 18:06:58 +02:00
José Valim 8cfaa07c41 Split on \r\n and \n accordingly during unicode compilation, closes #10991 2021-05-20 12:12:42 +02:00
Eksperimental d7cf95df0c Use proper Unicode reference (#10924) 2021-04-19 17:10:44 +02:00
Enrico Rivarola 589f9a6ce7 Sync unicode data to Unicode 13.0.0 (#10866) 2021-04-03 15:20:05 +02:00
Eksperimental 3ba7fe7aae Update Unicode upgrade instructions (#10862) 2021-04-01 22:11:23 +02:00
Eksperimental 92ec16c77a Update Unicode to version 13.0.0 (#10861) 2021-04-01 20:35:52 +02:00
Cẩm Huỳnh 5b59ce40da Optimize String.codepoints (#10633) 2021-01-08 19:20:27 +01:00
José Valim 0820b21a0d Consider CRLF on Windows, closes #10470 2020-11-01 17:25:06 +01:00
José Valim 255c09e452 Only check the mode after matching on the codepoint
This guarantees the compiler will not emit the code
that is first always checking the mode.
2020-10-03 10:20:08 +02:00
Alt D. Soy 7759bcccae Add :turkic mode option to String case functions (#10376)
This PR adds support of Turkic languages casing of the letter i according
to Unicode SpecialCasing.txt.
2020-10-03 09:52:18 +02:00
Eksperimental 0acab05f53 Use titlecase in Elixir, Erlang, EEx, Unicode, and Dialyzer (#9557) 2019-11-19 09:10:29 +01:00
José Valim f410cb13d1 Update to Unicode 12.1 2019-07-05 11:04:08 +02:00
José Valim 45ae6faeb5 Improve docs for built-in data types 2019-07-05 11:04:08 +02:00