Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Lexical Structure

This page describes the conversion from Larol source ASCII characters into tokens.

Larol source ASCII characters are converted into lexical tokens through lexical analysis. Tokens are constructed by matching consecutive input characters against the token productions shown below. Each source character belongs to at most one token, and each token consists of one or more consecutive source characters.

Discarded Characters

Whitespace and comments have no semantic meaning and are only required where they are necessary to separate adjacent tokens, which would otherwise be greedily lexed.

Whitespace

The following ASCII characters are considered whitespace and do not directly produce tokens.

  1. Character Tabulation,
  2. Line Feed,
  3. Line Tabulation,
  4. Form Feed,
  5. Carriage Return,
  6. Space.

Comments

Characters following, and including // are ignored until the next 'Line Feed' or 'Carriage Return' character.

Characters following, and including /* are ignored until after the occurrence of */.

Disambiguations

In cases where multiple different valid tokens match, disambiguation rules are used:

  1. tokens with a longer total length (in characters) are given higher precedence;
  2. all other tokens are given higher precedence than the identifier token.

Additional disambiguation rules should not be required for any inputs.

Unexpected Tokens

Any source character which:

  1. is not discarded, and
  2. does not belong to a token matched by the lexical token productions,

causes a lexical error.

A compliant implementation may recover from a lexical error and continue lexical analysis, producing additional tokens; alternatively, such an error could be fatal.

Syntax

Lexical Helper Productions

The following productions are helpers used in the token productions; they are not lexed into individual tokens.

letter = "A" | "B" | "C" | "D" | "E" | "F" | "G"
       | "H" | "I" | "J" | "K" | "L" | "M" | "N"
       | "O" | "P" | "Q" | "R" | "S" | "T" | "U"
       | "V" | "W" | "X" | "Y" | "Z"
       | "a" | "b" | "c" | "d" | "e" | "f" | "g"
       | "h" | "i" | "j" | "k" | "l" | "m" | "n"
       | "o" | "p" | "q" | "r" | "s" | "t" | "u"
       | "v" | "w" | "x" | "y" | "z" ;

digit = "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9" ;

hex_digit = digit | "A" | "B" | "C" | "D" | "E" | "F"
          | "a" | "b" | "c" | "d" | "e" | "f" ;

character = ? any ASCII source character ? ;

quotation_mark  = '"' ;
apostrophe      = "'" ;
reverse_solidus = "\" ;

escape_sequence = reverse_solidus, ( "n" | "r" | "t" | reverse_solidus | quotation_mark
                  | "x", hex_digit, hex_digit ) ;

char_escape_sequence = reverse_solidus, ( "n" | "r" | "t" | reverse_solidus | apostrophe
                  | "x", hex_digit, hex_digit ) ;

string_character = character - ( quotation_mark | reverse_solidus | ? line feed ? | ? carriage return ? ) ;

char_character = character - ( apostrophe | reverse_solidus | ? line feed ? | ? carriage return ? ) ;

exponent_part = "e", [ "-" ], digit, { [ "_" ], digit } ;

Lexical Token Productions

Each of the following productions represents an individual token.

identifier = ( letter | "_" ), { letter | digit | "_" } ;

(* Literals *)
integer_literal = digit, { [ "_" ], digit }
                | "0x", hex_digit, { [ "_" ], hex_digit } ;

float_literal = digit, { [ "_" ], digit }, ".", digit, { [ "_" ], digit }, [ exponent_part ] ;

boolean_literal = true | false ;

string_literal = quotation_mark, { string_character | escape_sequence }, quotation_mark ;

char_literal = apostrophe, ( char_character | char_escape_sequence ), apostrophe ;

(* Reserved Keywords *)
alloc     = "alloc"     ;
collect   = "collect"   ;
i8        = "i8"        ;
i16       = "i16"       ;
i32       = "i32"       ;
i64       = "i64"       ;
u8        = "u8"        ;
u16       = "u16"       ;
u32       = "u32"       ;
u64       = "u64"       ;
f32       = "f32"       ;
f64       = "f64"       ;
bool      = "bool"      ;
function  = "function"  ;
enum      = "enum"      ;
namespace = "namespace" ;
over      = "over"      ;
structure = "structure" ;
define    = "define"    ;
let       = "let"       ;
match     = "match"     ;
false     = "false"     ;
true      = "true"      ;
type      = "type"      ;

(* Symbols *)
ampersand            = "&" ;
ampersand_ampersand  = "&&" ;
bang                 = "!" ;
bang_equal           = "!=" ;
colon                = ":" ;
colon_colon          = "::" ;
comma                = "," ;
dot                  = "." ;
equal                = "=" ;
equal_equal          = "==" ;
greater_than         = ">" ;
greater_than_equal   = ">=" ;
left_brace           = "{" ;
left_bracket         = "[" ;
left_paren           = "(" ;
less_than            = "<" ;
less_than_equal      = "<=" ;
minus                = "-" ;
number_sign          = "#" ;
percent              = "%" ;
plus                 = "+" ;
question_mark        = "?" ;
right_brace          = "}" ;
right_bracket        = "]" ;
right_paren          = ")" ;
semicolon            = ";" ;
solidus              = "/" ;
star                 = "*" ;
vertical_bar_vertical_bar = "||" ;