Commit Graph
24 Commits
Author SHA1 Message Date
José Valim ec50a88ee4 Validate normalized tokenizer identifiers
Ensure identifiers that require normalization are checked again
after normalization so decomposed source cannot normalize into
unsupported codepoints.

Closes #15419.
2026-06-18 10:55:09 +02:00
José Valim 175f54869b Update Unicode to version 17.0.0 (#14760)
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance

Please read the release notes by visiting
<http://www.unicode.org/versions/Unicode17.0.0/>.
2025-09-10 09:09:44 +02:00
Jonatan Männchen 980500551c Inline Licence Info & Copyright for non-test files (#14256) 2025-02-07 21:11:09 +01:00
José Valim c08a9b371c Update CHANGELOG 2024-07-02 15:21:25 +02:00
José Valim f01dc17d85 Optimize reversing logic
Rule #1:

    Enum.reverse([1, 2, 3]) ++ [4, 5, 6]

is equivalent but slower than

    Enum.reverse([1, 2, 3], [4, 5, 6])

Rule #2:

    Enum.reverse([1, 2, 3] ++ [4, 5, 6])

is equivalent but slower than:

    Enum.reverse([4, 5, 6], Enum.reverse([1, 2, 3]))

Rule #3

    Enum.reverse(Enum.reverse([1, 2, 3]))

is the same as:

    [1, 2, 3]
2024-07-02 14:51:49 +02:00
José Valim 9924afff5d Remove highly restrictive scriptset support 2024-07-02 14:41:57 +02:00
Luc Fueston c83334e5f2 Directional confusability protection & allowing script mixing in identifiers when separated by underscores (#13693)
* update support for uts39 from unicode 15

* follow uts39's recco that it's not necessary to require
  idents to be single-script (they call out proglang idents,
  reference the new uts55-5). We use a heuristic derived from
  the concept of identifier chunks from uts55-5, to allow
  idents like foo_bar_baz where each chunk around the _ can be
  single-or-highly-restrictive

* provide directional confusability detection, by reversing
  spans of direction-changed chars in idents for bidi_skeleton,
  see issue #12929
2024-07-02 14:07:20 +02:00
José Valim 47079bc0b2 Add more deprecations scheduled to v1.18 2024-05-26 21:44:10 +02:00
Andrea Leopardi dd295f4bc7 Fix spelling of "hard-coded" 2023-07-09 16:08:12 +02:00
sabiwara 29bf449f3f Evaluate formating charlists as ~c sigils (#12064) 2022-08-10 06:47:00 +02:00
José Valim fbf94613f3 Fix regression on identifier\\ being interpreted as a single identifier 2022-06-30 15:52:18 +02:00
José Valim 850c7ddfe4 Normalize unicode characters once in tokenizer 2022-05-23 12:14:39 +03:00
José Valim 25212f2db1 Do not attempt to suggest starting tokens
It is hard to suggest because it is not possible to
know if the user intended an operator, alias, or
identifier, so we can give incorrect feedback.
2022-05-23 12:14:39 +03:00
Luc Fueston 31cbdbd6e0 nfc and additional normalizations for identifiers (#11859)
* nfc by default
* suggest nfkc on unexpected tokens when nfkc would have worked
* support additional normalizations (just micro/mu for now)
* parser, formatter tests
* tweak Macro.inner_classify(atom) not to rely on not_nfc causing error
2022-05-23 09:44:44 +02:00
José Valim f067abaf11 Remove duplication in unicode tokenizer 2022-04-09 09:46:56 +02:00
José Valim 2676f8b221 Allow only single- and highly restricted mixed-scripts in identifiers (UTS39 C3) (#11621) 2022-02-11 20:44:06 +01:00
Luc Fueston 1b80c8c838 Don't allow restricted characters in identifiers (UTS 39, C1) (#11580) 2022-01-17 19:42:36 +01:00
José Valim 8cfaa07c41 Split on \r\n and \n accordingly during unicode compilation, closes #10991 2021-05-20 12:12:42 +02:00
Enrico Rivarola 589f9a6ce7 Sync unicode data to Unicode 13.0.0 (#10866) 2021-04-03 15:20:05 +02:00
José Valim 0820b21a0d Consider CRLF on Windows, closes #10470 2020-11-01 17:25:06 +01:00
Eric Meadows-Jönsson 1f974f869e Move unreachable function check from xref to group pass (#9168) 2019-06-28 10:25:49 +02:00
Frank McGeough 24acb605ba Run formatter on lib/elixir/unicode/tokenizer.ex (#6813) 2017-10-10 18:10:35 +02:00
José Valim 19b1668cfd Do not rely on on_load because it breaks releases 2017-05-29 15:50:06 +02:00
José Valim 2f663ce3a2 Add non-quoted Unicode atoms and variables (#6158)
It follows Unicode Annex #31.

See the Unicode Syntax document for a more in-depth reference.
2017-05-27 22:43:05 +02:00