ANTLR4 to TextMate grammar translation
Here is an example of what would happen if we translated the given ANTLR4 lexer grammar naively into a TextMate syntax-highlighting grammar:
We call this the shadowing problem. The core issue is that ANTLR-generated lexers differ in behavior from TextMate-based syntax highlighters when the input matches multiple rules:
- an ANTLR lexer picks the rule that matches the most characters from the input (→ longest-match)
- a TextMate tokenizer picks the first rule that finds any match (→ leftmost-match)
This also applies to alternatives (e.g. ('a' | 'ab')) within a single rule due to the way
the Oniguruma regex engine works.
Just reordering the rules doesn’t help:
// TODO: write about the solution