Lexical Structure
This page describes the conversion from Larol source ASCII characters into tokens.
Larol source ASCII characters are converted into lexical tokens through lexical analysis. Tokens are constructed by matching consecutive input characters against the token productions shown below. Each source character belongs to at most one token, and each token consists of one or more consecutive source characters.
Discarded Characters
Whitespace and comments have no semantic meaning and are only required where they are necessary to separate adjacent tokens, which would otherwise be greedily lexed.
Whitespace
The following ASCII characters are considered whitespace and do not directly produce tokens.
- Character Tabulation,
- Line Feed,
- Line Tabulation,
- Form Feed,
- Carriage Return,
- Space.
Comments
Characters following, and including // are ignored until the next 'Line Feed' or 'Carriage Return' character.
Characters following, and including /* are ignored until after the occurrence of */.
Disambiguations
In cases where multiple different valid tokens match, disambiguation rules are used:
- tokens with a longer total length (in characters) are given higher precedence;
- all other tokens are given higher precedence than the
identifiertoken.
Additional disambiguation rules should not be required for any inputs.
Unexpected Tokens
Any source character which:
- is not discarded, and
- does not belong to a token matched by the lexical token productions,
causes a lexical error.
A compliant implementation may recover from a lexical error and continue lexical analysis, producing additional tokens; alternatively, such an error could be fatal.
Syntax
Lexical Helper Productions
The following productions are helpers used in the token productions; they are not lexed into individual tokens.
letter = "A" | "B" | "C" | "D" | "E" | "F" | "G"
| "H" | "I" | "J" | "K" | "L" | "M" | "N"
| "O" | "P" | "Q" | "R" | "S" | "T" | "U"
| "V" | "W" | "X" | "Y" | "Z"
| "a" | "b" | "c" | "d" | "e" | "f" | "g"
| "h" | "i" | "j" | "k" | "l" | "m" | "n"
| "o" | "p" | "q" | "r" | "s" | "t" | "u"
| "v" | "w" | "x" | "y" | "z" ;
digit = "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9" ;
hex_digit = digit | "A" | "B" | "C" | "D" | "E" | "F"
| "a" | "b" | "c" | "d" | "e" | "f" ;
character = ? any ASCII source character ? ;
quotation_mark = '"' ;
apostrophe = "'" ;
reverse_solidus = "\" ;
escape_sequence = reverse_solidus, ( "n" | "r" | "t" | reverse_solidus | quotation_mark
| "x", hex_digit, hex_digit ) ;
char_escape_sequence = reverse_solidus, ( "n" | "r" | "t" | reverse_solidus | apostrophe
| "x", hex_digit, hex_digit ) ;
string_character = character - ( quotation_mark | reverse_solidus | ? line feed ? | ? carriage return ? ) ;
char_character = character - ( apostrophe | reverse_solidus | ? line feed ? | ? carriage return ? ) ;
exponent_part = "e", [ "-" ], digit, { [ "_" ], digit } ;
Lexical Token Productions
Each of the following productions represents an individual token.
identifier = ( letter | "_" ), { letter | digit | "_" } ;
(* Literals *)
integer_literal = digit, { [ "_" ], digit }
| "0x", hex_digit, { [ "_" ], hex_digit } ;
float_literal = digit, { [ "_" ], digit }, ".", digit, { [ "_" ], digit }, [ exponent_part ] ;
boolean_literal = true | false ;
string_literal = quotation_mark, { string_character | escape_sequence }, quotation_mark ;
char_literal = apostrophe, ( char_character | char_escape_sequence ), apostrophe ;
(* Reserved Keywords *)
alloc = "alloc" ;
collect = "collect" ;
i8 = "i8" ;
i16 = "i16" ;
i32 = "i32" ;
i64 = "i64" ;
u8 = "u8" ;
u16 = "u16" ;
u32 = "u32" ;
u64 = "u64" ;
f32 = "f32" ;
f64 = "f64" ;
bool = "bool" ;
function = "function" ;
enum = "enum" ;
namespace = "namespace" ;
over = "over" ;
structure = "structure" ;
define = "define" ;
let = "let" ;
match = "match" ;
false = "false" ;
true = "true" ;
type = "type" ;
(* Symbols *)
ampersand = "&" ;
ampersand_ampersand = "&&" ;
bang = "!" ;
bang_equal = "!=" ;
colon = ":" ;
colon_colon = "::" ;
comma = "," ;
dot = "." ;
equal = "=" ;
equal_equal = "==" ;
greater_than = ">" ;
greater_than_equal = ">=" ;
left_brace = "{" ;
left_bracket = "[" ;
left_paren = "(" ;
less_than = "<" ;
less_than_equal = "<=" ;
minus = "-" ;
number_sign = "#" ;
percent = "%" ;
plus = "+" ;
question_mark = "?" ;
right_brace = "}" ;
right_bracket = "]" ;
right_paren = ")" ;
semicolon = ";" ;
solidus = "/" ;
star = "*" ;
vertical_bar_vertical_bar = "||" ;