Commit Graph
113 Commits
Author SHA1 Message Date
José Valim 51884ca511 Fix regression on identifier\\ being interpreted as a single identifier 2022-06-30 15:52:18 +02:00
José Valim 5deafbdc89 Normalize unicode characters once in tokenizer 2022-05-23 12:14:39 +03:00
José Valim f431aefcac Do not attempt to suggest starting tokens
It is hard to suggest because it is not possible to
know if the user intended an operator, alias, or
identifier, so we can give incorrect feedback.
2022-05-23 12:14:39 +03:00
Luc Fueston e7001455d7 nfc and additional normalizations for identifiers (#11859)
* nfc by default
* suggest nfkc on unexpected tokens when nfkc would have worked
* support additional normalizations (just micro/mu for now)
* parser, formatter tests
* tweak Macro.inner_classify(atom) not to rely on not_nfc causing error
2022-05-23 09:44:44 +02:00
John Bampton 04554ba4e5 Remove trailing whitespace (#11782) 2022-04-25 15:25:31 +02:00
José Valim e8b24cd3f0 Remove duplication in unicode tokenizer 2022-04-09 09:46:56 +02:00
José Valim 18c44e2e4f Simplify process of copying SpecialCasing.txt 2022-02-13 09:01:05 +01:00
Eksperimental 0a07484d83 Improve Unicode module update instructions (#11622) 2022-02-13 08:54:24 +01:00
José Valim 9e0ae6140a Allow only single- and highly restricted mixed-scripts in identifiers (UTS39 C3) (#11621) 2022-02-11 20:44:06 +01:00
José Valim 826d2c8679 Do not emit warnings when formatting code
Closes #11595.
2022-01-26 13:37:53 +01:00
José Valim b401d15296 Run unicode security linting on more tokens 2022-01-21 10:50:02 +01:00
Luc Fueston a4689262d8 Warn on confusable non-ascii identifiers (UTS 39, C2) (#11582) 2022-01-21 09:37:25 +01:00
Luc Fueston 4a5fcbe04a Don't allow restricted characters in identifiers (UTS 39, C1) (#11580) 2022-01-17 19:42:36 +01:00
José Valim 5a2b24d21f Group Unicode upcase/downcase by prefix (#11310)
The patch computes byte lookups based on the prefix. For example,
Á, É, etc all have the same prefix <<195>>, so they are lumped
together for lookup and then we just do a byte lookup later. We
tried doing the byte lookup on 64-element tuple (since the byte
is always within 0b10000000 and 0b10111111) but that's slower,
especially because we need to check the byte range for invalid
Unicode, so instead the last byte lookup is a case. Grouping the
top-level lookup makes the cost of a miss 3x cheaper albeit a
hit is 10% more expensive and reduces bytecode size.
2021-10-14 14:44:13 +02:00
Eksperimental d962ddb5b0 Upgrade Unicode database to v14.0.0 (#11308)
Closes #11288.
2021-10-12 18:05:10 +02:00
José Valim 3f5b3f00f1 Depend on Erlang for computing grapheme clusters (#11024)
Erlang ships with its own embedding of the Unicode
Codebase for a couple releases and this release changes
Elixir to depend on it in order to compute grapheme
clusters.

The Erlang implementation was up to 2x faster in low
codepoints (such as latin1) while the Elixir one could
be faster up to 3x in high codepoints (such as emoji)
so at the end the performance results are roughly the
same. As a benefit, we no longer need to ship our copy
of the grapheme cluster algorithm, which would take up
to 250kB in disk and more than 15 seconds to compile.

Note we still keep our own String downcase and upcase
algorithms, as our version is considerably more efficient
on all cases since it works exclusively with binaries
(more than 5x faster).
2021-06-01 12:44:55 +02:00
José Valim 0123f8680f Reduce size of Unicode module in 33% by not duplicating codepoints 2021-05-25 18:06:58 +02:00
José Valim 151f11e0b8 Split on \r\n and \n accordingly during unicode compilation, closes #10991 2021-05-20 12:12:42 +02:00
Eksperimental c1e622950c Use proper Unicode reference (#10924) 2021-04-19 17:10:44 +02:00
Enrico Rivarola 218c35a041 Sync unicode data to Unicode 13.0.0 (#10866) 2021-04-03 15:20:05 +02:00
Eksperimental 0beb4f42fe Update Unicode upgrade instructions (#10862) 2021-04-01 22:11:23 +02:00
Eksperimental 37a09feaaa Update Unicode to version 13.0.0 (#10861) 2021-04-01 20:35:52 +02:00
Cẩm Huỳnh ef16468175 Optimize String.codepoints (#10633) 2021-01-08 19:20:27 +01:00
José Valim c36791091f Consider CRLF on Windows, closes #10470 2020-11-01 17:25:06 +01:00
José Valim a1c90e6a89 Only check the mode after matching on the codepoint
This guarantees the compiler will not emit the code
that is first always checking the mode.
2020-10-03 10:20:08 +02:00
Alt D. Soy 8ff465fa49 Add :turkic mode option to String case functions (#10376)
This PR adds support of Turkic languages casing of the letter i according
to Unicode SpecialCasing.txt.
2020-10-03 09:52:18 +02:00
Eksperimental b0d4288174 Use titlecase in Elixir, Erlang, EEx, Unicode, and Dialyzer (#9557) 2019-11-19 09:10:29 +01:00
José Valim bffb79e553 Update to Unicode 12.1 2019-07-05 11:04:08 +02:00
José Valim 0e6b8732ba Improve docs for built-in data types 2019-07-05 11:04:08 +02:00
José Valim 076dece778 Avoid unicode codepoint computation for lower bytes 2019-07-04 09:38:55 +02:00
Eric Meadows-Jönsson 9b6ec7bae7 Move unreachable function check from xref to group pass (#9168) 2019-06-28 10:25:49 +02:00
Eksperimental d4f8d419e4 Use https instead of http when servers are redirecting (#9094)
The servers are automatically redirecting from http:// to https://
2019-05-29 07:37:54 +02:00
José Valim 93fb75d793 Optimize Unicode graphemes
We only check for Pictographic once we find a ZWJ.
2019-03-27 01:46:34 +01:00
Eksperimental 99d919e0c8 Use "code point" instead of "codepoint" (#8774)
That's the term it is used in the Unicode standard.
2019-02-06 11:21:03 +01:00
Xavier Noria ee5a7fa0a4 Clarify String.next_codepoint/1 re invalid UTF-8 (#8489) 2018-12-08 16:41:57 +01:00
Sun Yaozhu 1967a06abe Fix ZWJ handling in Unicode grapheme clusters (#8361) 2018-11-04 19:56:36 +01:00
José Valim 67b7ffce61 Print number of tests from graphemes tests 2018-10-02 12:25:50 +02:00
José Valim 54cb02c240 Rely on Erlang/OTP normalization which is faster 2018-07-12 13:28:58 +02:00
José Valim c995ba968f Update to Unicode 11 2018-06-06 11:22:36 +02:00
Michał Muskała 722aef3791 Use binary_part/3 instead of :binary.part/3
binary_part is a guard BIF and is slightly more efficient to use then
:binary.part which does a fully qualified external call.
2018-04-30 13:22:19 +02:00
José Valim a7eb9adfe9 Consider case ignorable characters on Greek downcasing, closes #7149
Note this does not impact the runtime cost of other downcasing
operations. The final beam file grew only in 8kb.
2018-01-09 13:59:13 +01:00
José Valim 44a8f665b1 Improvements to greek handling of downcase, closes #7149 2017-12-26 23:11:17 +01:00
José Valim 6ecfbc546c Optimize String.downcase/1
Thanks to @michalmuskala for feedback and review.
2017-12-20 12:58:17 +01:00
José Valim 18288881c2 Provide :greek and :ascii mappings in String upcase/downcase/capitalize
Closes #7105.
2017-12-20 11:54:06 +01:00
José Valim b4aac2d7be Do not indent right-side of pipelines in the formatter (#7103) 2017-12-11 21:07:30 +01:00
José Valim b3992aa574 Minimize generated beam size by using ranges
This approach is less performant when processing
sigma but compiles twice faster and takes two times
less space in disk.
2017-10-30 18:45:16 +01:00
Ben Olive dfcc88fc33 Handle final sigma unicode exception
Signed-off-by: José Valim <jose.valim@plataformatec.com.br>
2017-10-30 18:45:16 +01:00
José Valim c217203558 Improvements to previously formatted files 2017-10-11 22:59:30 +02:00
José Valim 97ad07b7ab Make sure to not break guards on clauses with comments 2017-10-11 21:25:21 +02:00
José Valim 660838b399 Format more files with little formatting requirement 2017-10-11 00:28:10 +02:00