Kernel.raise messages should not start with uppercase, and should not
have trailing punctuation.
Mix.raise messages should start with uppercase, and should not have a trailing
puntuation.
# starting with uppercase
# exclude (Mix.raise, and any message that starts with all uppecase such as: I or IO
ag '(?<!Mix\.)raise\s+\"+(?!([A-Z]+\b))[A-Z]'
# Mix.raise should start with upper case
# except when `mix` or `rebar` are mentioned
ag 'Mix\.raise\s+\"(?!(mix|rebar))+[a-z]'
# ending in period
ag "raise\s+\"+.+\.\"\s*\n"
Many functions relied on next_grapheme, which would traverse
the whole string generating a lot of graphemes as garbage along
the way.
This commit introduces next_grapheme_size, which returns the
size in bytes of the next grapheme, avoiding the generation
of the garbage and optimizing functions that need just the
byte_size.
Functions optimized are: String.at, String.split_at, String.slice
and String.length.
This new implementation is no longer linear without affecting
smaller samples. For a string that is 100 bytes long, it is
25x faster than the previous implementation.
Compiling in parallel only bought us 2s (out of 24s) since
compiling graphemes is what takes most of the time but it
considerably increases memory usage, by about 30%, since both
compilation targets are loaded into memory at once.
Based on a suggestion @krestenkrab and an initial implementation
proposed in #1902, we can reduce the total size of the String.Graphemes
beam file by ensuring the function clauses for next_grapheme are
similar.
We have done this by implementing next_grapheme in a tail recursive
fashion and using binary_part to select the first bits. The end result
is a consistently 33% faster graphemes operation for ascii strings
and only 10% slower operation for heavy unicode strings (which may
be further optimazable).