When you are working on large files and you add new code,
you may forget to add a `do` or an `end`. In those cases,
the error message usually points to the first `do` or the
last `end` in the file, which are usually far away from
the source of the error.
This pull request adds a simple heuristic based on the
indentation of the tokens, to try to provide hints of
where the source may be. Those hints are not deterministic
but they may be able to point users to the source of the
problem.
For example, in this case:
defmodule MyApp do
def one do
# end
def two do
end
end
we know that we now that `def two do` is happening on the
same indentation as `def one do`, which may mean that
`def one do` was not closed properly. We store this as a
hint in case the terminators do not match later.
Similarly, in the case below:
defmodule MyApp do
def one
end
def two do
end
end
The `end` on line 3 will end-up closing the defmodule `do`,
on line 1. Because their indentation do not match, it may
be that there is a missing `do`, where the `end` was supposed
to align.
Some basic testing show those heuristics work on the majority
of the cases, but we will only be sure when we have enough
feedback from the community.
Right now, tokens are "{Token, Location}" or "{Token, Location, Value}".
This commit changes "Location" from "{Line, StartColumn, EndColumn}" to
"{Line, {StartColumn, EndColumn}, Meta}" where "Meta" can be anything.
This will be used for things such as storing the format of integers.
Standardizes the use of comma leaving a white space after it whenever applicable.
Note: It does not enforce this in quantifiers in regular expressions such as in: `x{1,3}`
Before this commit, we handled char literals (like `?a`) in the
tokenizer, turning a literal like `?a` into the token `{:number, _,
97}` (thus indistinguishable from the literal `97` at the parsing
stage). This led to error messages with the integer for the character
instead of the character literal, e.g.:
iex> :ok ?a
** (SyntaxError) iex:11: syntax error before: 97
With this commit, we now turn `?a` into the token `{:char, _,
97}` (which is the same token used by Erlang for Erlang char literals
like `$a`); since it's the same token as in Erlang, the parser will now
output the char literal as an Erlang char (`?a` would be printed as
`$a`). We hijack the error message in elixir_errors.erl to end up with
the correct message:
iex> :ok ?a
** (SyntaxError) iex:11: syntax error before: ?a
Add column info for each token in elixir_tokenizer.
Change the format of location info from `Line` to `[Line, BeginColumn,
EndColumn]`. Pass the current column after the current line in
`elixir_tokenizer:tokenize`. Reflect the change in related modules.
This refactoring helps us clean up the interpolation code and
guarantee it works as expected since it goes under the same
tokenization rules. It also helps provide better sigil handling
(going to be improved in upcoming commits).
We also rename Binary.Chars to String.Chars.
As Binary.Chars was the protocol used for interpolation,
it is semantically correct to have a protocol that works
on strings, instead of raw binaries.
This means the second item in the AST is no longer an integer
(representing the line), but a keywords list. Code that relies
on the line information from AST or that manually generate AST
nodes need to be properly updated.