José Valim
51884ca511
Fix regression on identifier\\ being interpreted as a single identifier
2022-06-30 15:52:18 +02:00
José Valim
5deafbdc89
Normalize unicode characters once in tokenizer
2022-05-23 12:14:39 +03:00
José Valim
f431aefcac
Do not attempt to suggest starting tokens
...
It is hard to suggest because it is not possible to
know if the user intended an operator, alias, or
identifier, so we can give incorrect feedback.
2022-05-23 12:14:39 +03:00
Luc Fueston
e7001455d7
nfc and additional normalizations for identifiers ( #11859 )
...
* nfc by default
* suggest nfkc on unexpected tokens when nfkc would have worked
* support additional normalizations (just micro/mu for now)
* parser, formatter tests
* tweak Macro.inner_classify(atom) not to rely on not_nfc causing error
2022-05-23 09:44:44 +02:00
John Bampton
04554ba4e5
Remove trailing whitespace ( #11782 )
2022-04-25 15:25:31 +02:00
José Valim
e8b24cd3f0
Remove duplication in unicode tokenizer
2022-04-09 09:46:56 +02:00
José Valim
18c44e2e4f
Simplify process of copying SpecialCasing.txt
2022-02-13 09:01:05 +01:00
Eksperimental
0a07484d83
Improve Unicode module update instructions ( #11622 )
2022-02-13 08:54:24 +01:00
José Valim
9e0ae6140a
Allow only single- and highly restricted mixed-scripts in identifiers (UTS39 C3) ( #11621 )
2022-02-11 20:44:06 +01:00
José Valim
826d2c8679
Do not emit warnings when formatting code
...
Closes #11595 .
2022-01-26 13:37:53 +01:00
José Valim
b401d15296
Run unicode security linting on more tokens
2022-01-21 10:50:02 +01:00
Luc Fueston
a4689262d8
Warn on confusable non-ascii identifiers (UTS 39, C2) ( #11582 )
2022-01-21 09:37:25 +01:00
Luc Fueston
4a5fcbe04a
Don't allow restricted characters in identifiers (UTS 39, C1) ( #11580 )
2022-01-17 19:42:36 +01:00
José Valim
5a2b24d21f
Group Unicode upcase/downcase by prefix ( #11310 )
...
The patch computes byte lookups based on the prefix. For example,
Á, É, etc all have the same prefix <<195>>, so they are lumped
together for lookup and then we just do a byte lookup later. We
tried doing the byte lookup on 64-element tuple (since the byte
is always within 0b10000000 and 0b10111111) but that's slower,
especially because we need to check the byte range for invalid
Unicode, so instead the last byte lookup is a case. Grouping the
top-level lookup makes the cost of a miss 3x cheaper albeit a
hit is 10% more expensive and reduces bytecode size.
2021-10-14 14:44:13 +02:00
Eksperimental
d962ddb5b0
Upgrade Unicode database to v14.0.0 ( #11308 )
...
Closes #11288 .
2021-10-12 18:05:10 +02:00
José Valim
3f5b3f00f1
Depend on Erlang for computing grapheme clusters ( #11024 )
...
Erlang ships with its own embedding of the Unicode
Codebase for a couple releases and this release changes
Elixir to depend on it in order to compute grapheme
clusters.
The Erlang implementation was up to 2x faster in low
codepoints (such as latin1) while the Elixir one could
be faster up to 3x in high codepoints (such as emoji)
so at the end the performance results are roughly the
same. As a benefit, we no longer need to ship our copy
of the grapheme cluster algorithm, which would take up
to 250kB in disk and more than 15 seconds to compile.
Note we still keep our own String downcase and upcase
algorithms, as our version is considerably more efficient
on all cases since it works exclusively with binaries
(more than 5x faster).
2021-06-01 12:44:55 +02:00
José Valim
0123f8680f
Reduce size of Unicode module in 33% by not duplicating codepoints
2021-05-25 18:06:58 +02:00
José Valim
151f11e0b8
Split on \r\n and \n accordingly during unicode compilation, closes #10991
2021-05-20 12:12:42 +02:00
Eksperimental
c1e622950c
Use proper Unicode reference ( #10924 )
2021-04-19 17:10:44 +02:00
Enrico Rivarola
218c35a041
Sync unicode data to Unicode 13.0.0 ( #10866 )
2021-04-03 15:20:05 +02:00
Eksperimental
0beb4f42fe
Update Unicode upgrade instructions ( #10862 )
2021-04-01 22:11:23 +02:00
Eksperimental
37a09feaaa
Update Unicode to version 13.0.0 ( #10861 )
2021-04-01 20:35:52 +02:00
Cẩm Huỳnh
ef16468175
Optimize String.codepoints ( #10633 )
2021-01-08 19:20:27 +01:00
José Valim
c36791091f
Consider CRLF on Windows, closes #10470
2020-11-01 17:25:06 +01:00
José Valim
a1c90e6a89
Only check the mode after matching on the codepoint
...
This guarantees the compiler will not emit the code
that is first always checking the mode.
2020-10-03 10:20:08 +02:00
Alt D. Soy
8ff465fa49
Add :turkic mode option to String case functions ( #10376 )
...
This PR adds support of Turkic languages casing of the letter i according
to Unicode SpecialCasing.txt.
2020-10-03 09:52:18 +02:00
Eksperimental
b0d4288174
Use titlecase in Elixir, Erlang, EEx, Unicode, and Dialyzer ( #9557 )
2019-11-19 09:10:29 +01:00
José Valim
bffb79e553
Update to Unicode 12.1
2019-07-05 11:04:08 +02:00
José Valim
0e6b8732ba
Improve docs for built-in data types
2019-07-05 11:04:08 +02:00
José Valim
076dece778
Avoid unicode codepoint computation for lower bytes
2019-07-04 09:38:55 +02:00
Eric Meadows-Jönsson
9b6ec7bae7
Move unreachable function check from xref to group pass ( #9168 )
2019-06-28 10:25:49 +02:00
Eksperimental
d4f8d419e4
Use https instead of http when servers are redirecting ( #9094 )
...
The servers are automatically redirecting from http:// to https://
2019-05-29 07:37:54 +02:00
José Valim
93fb75d793
Optimize Unicode graphemes
...
We only check for Pictographic once we find a ZWJ.
2019-03-27 01:46:34 +01:00
Eksperimental
99d919e0c8
Use "code point" instead of "codepoint" ( #8774 )
...
That's the term it is used in the Unicode standard.
2019-02-06 11:21:03 +01:00
Xavier Noria
ee5a7fa0a4
Clarify String.next_codepoint/1 re invalid UTF-8 ( #8489 )
2018-12-08 16:41:57 +01:00
Sun Yaozhu
1967a06abe
Fix ZWJ handling in Unicode grapheme clusters ( #8361 )
2018-11-04 19:56:36 +01:00
José Valim
67b7ffce61
Print number of tests from graphemes tests
2018-10-02 12:25:50 +02:00
José Valim
54cb02c240
Rely on Erlang/OTP normalization which is faster
2018-07-12 13:28:58 +02:00
José Valim
c995ba968f
Update to Unicode 11
2018-06-06 11:22:36 +02:00
Michał Muskała
722aef3791
Use binary_part/3 instead of :binary.part/3
...
binary_part is a guard BIF and is slightly more efficient to use then
:binary.part which does a fully qualified external call.
2018-04-30 13:22:19 +02:00
José Valim
a7eb9adfe9
Consider case ignorable characters on Greek downcasing, closes #7149
...
Note this does not impact the runtime cost of other downcasing
operations. The final beam file grew only in 8kb.
2018-01-09 13:59:13 +01:00
José Valim
44a8f665b1
Improvements to greek handling of downcase, closes #7149
2017-12-26 23:11:17 +01:00
José Valim
6ecfbc546c
Optimize String.downcase/1
...
Thanks to @michalmuskala for feedback and review.
2017-12-20 12:58:17 +01:00
José Valim
18288881c2
Provide :greek and :ascii mappings in String upcase/downcase/capitalize
...
Closes #7105 .
2017-12-20 11:54:06 +01:00
José Valim
b4aac2d7be
Do not indent right-side of pipelines in the formatter ( #7103 )
2017-12-11 21:07:30 +01:00
José Valim
b3992aa574
Minimize generated beam size by using ranges
...
This approach is less performant when processing
sigma but compiles twice faster and takes two times
less space in disk.
2017-10-30 18:45:16 +01:00
Ben Olive
dfcc88fc33
Handle final sigma unicode exception
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2017-10-30 18:45:16 +01:00
José Valim
c217203558
Improvements to previously formatted files
2017-10-11 22:59:30 +02:00
José Valim
97ad07b7ab
Make sure to not break guards on clauses with comments
2017-10-11 21:25:21 +02:00
José Valim
660838b399
Format more files with little formatting requirement
2017-10-11 00:28:10 +02:00