José Valim
ec50a88ee4
Validate normalized tokenizer identifiers
...
Ensure identifiers that require normalization are checked again
after normalization so decomposed source cannot normalize into
unsupported codepoints.
Closes #15419 .
2026-06-18 10:55:09 +02:00
José Valim
175f54869b
Update Unicode to version 17.0.0 ( #14760 )
...
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance
Please read the release notes by visiting
<http://www.unicode.org/versions/Unicode17.0.0/ >.
2025-09-10 09:09:44 +02:00
Jonatan Männchen
980500551c
Inline Licence Info & Copyright for non-test files ( #14256 )
2025-02-07 21:11:09 +01:00
José Valim
c08a9b371c
Update CHANGELOG
2024-07-02 15:21:25 +02:00
José Valim
f01dc17d85
Optimize reversing logic
...
Rule #1 :
Enum.reverse([1, 2, 3]) ++ [4, 5, 6]
is equivalent but slower than
Enum.reverse([1, 2, 3], [4, 5, 6])
Rule #2 :
Enum.reverse([1, 2, 3] ++ [4, 5, 6])
is equivalent but slower than:
Enum.reverse([4, 5, 6], Enum.reverse([1, 2, 3]))
Rule #3
Enum.reverse(Enum.reverse([1, 2, 3]))
is the same as:
[1, 2, 3]
2024-07-02 14:51:49 +02:00
José Valim
9924afff5d
Remove highly restrictive scriptset support
2024-07-02 14:41:57 +02:00
Luc Fueston
c83334e5f2
Directional confusability protection & allowing script mixing in identifiers when separated by underscores ( #13693 )
...
* update support for uts39 from unicode 15
* follow uts39's recco that it's not necessary to require
idents to be single-script (they call out proglang idents,
reference the new uts55-5). We use a heuristic derived from
the concept of identifier chunks from uts55-5, to allow
idents like foo_bar_baz where each chunk around the _ can be
single-or-highly-restrictive
* provide directional confusability detection, by reversing
spans of direction-changed chars in idents for bidi_skeleton,
see issue #12929
2024-07-02 14:07:20 +02:00
José Valim
47079bc0b2
Add more deprecations scheduled to v1.18
2024-05-26 21:44:10 +02:00
Andrea Leopardi
dd295f4bc7
Fix spelling of "hard-coded"
2023-07-09 16:08:12 +02:00
sabiwara
29bf449f3f
Evaluate formating charlists as ~c sigils ( #12064 )
2022-08-10 06:47:00 +02:00
José Valim
fbf94613f3
Fix regression on identifier\\ being interpreted as a single identifier
2022-06-30 15:52:18 +02:00
José Valim
850c7ddfe4
Normalize unicode characters once in tokenizer
2022-05-23 12:14:39 +03:00
José Valim
25212f2db1
Do not attempt to suggest starting tokens
...
It is hard to suggest because it is not possible to
know if the user intended an operator, alias, or
identifier, so we can give incorrect feedback.
2022-05-23 12:14:39 +03:00
Luc Fueston
31cbdbd6e0
nfc and additional normalizations for identifiers ( #11859 )
...
* nfc by default
* suggest nfkc on unexpected tokens when nfkc would have worked
* support additional normalizations (just micro/mu for now)
* parser, formatter tests
* tweak Macro.inner_classify(atom) not to rely on not_nfc causing error
2022-05-23 09:44:44 +02:00
José Valim
f067abaf11
Remove duplication in unicode tokenizer
2022-04-09 09:46:56 +02:00
José Valim
2676f8b221
Allow only single- and highly restricted mixed-scripts in identifiers (UTS39 C3) ( #11621 )
2022-02-11 20:44:06 +01:00
Luc Fueston
1b80c8c838
Don't allow restricted characters in identifiers (UTS 39, C1) ( #11580 )
2022-01-17 19:42:36 +01:00
José Valim
8cfaa07c41
Split on \r\n and \n accordingly during unicode compilation, closes #10991
2021-05-20 12:12:42 +02:00
Enrico Rivarola
589f9a6ce7
Sync unicode data to Unicode 13.0.0 ( #10866 )
2021-04-03 15:20:05 +02:00
José Valim
0820b21a0d
Consider CRLF on Windows, closes #10470
2020-11-01 17:25:06 +01:00
Eric Meadows-Jönsson
1f974f869e
Move unreachable function check from xref to group pass ( #9168 )
2019-06-28 10:25:49 +02:00
Frank McGeough
24acb605ba
Run formatter on lib/elixir/unicode/tokenizer.ex ( #6813 )
2017-10-10 18:10:35 +02:00
José Valim
19b1668cfd
Do not rely on on_load because it breaks releases
2017-05-29 15:50:06 +02:00
José Valim
2f663ce3a2
Add non-quoted Unicode atoms and variables ( #6158 )
...
It follows Unicode Annex #31 .
See the Unicode Syntax document for a more in-depth reference.
2017-05-27 22:43:05 +02:00