<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[SegniAT]]></title><description><![CDATA[I write on projects I build for fun.]]></description><link>https://segni.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6aaf9e9d355dadd7ed746dca/248c6d52-3d00-4a52-9c19-99c4ebe8a8f7.jpg</url><title>SegniAT</title><link>https://segni.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Tue, 29 Sep 2026 07:16:33 GMT</lastBuildDate><atom:link href="https://segni.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Building a development environment for Monkey: Part 1 - Syntax highlighting with Tree-sitter]]></title><description><![CDATA[Intro
A few years ago on a quest of learning Go, I followed Writing An Interpreter In Go and built Monkey, a small interpreted programming language. Recently, I came back to it with a different goal: ]]></description><link>https://segni.hashnode.dev/building-a-development-environment-for-monkey-part-1-syntax-highlighting-with-tree-sitter</link><guid isPermaLink="true">https://segni.hashnode.dev/building-a-development-environment-for-monkey-part-1-syntax-highlighting-with-tree-sitter</guid><category><![CDATA[interpreter]]></category><category><![CDATA[tree-sitter]]></category><category><![CDATA[programming languages]]></category><category><![CDATA[syntax highlighting]]></category><dc:creator><![CDATA[Segni Adeba]]></dc:creator><pubDate>Sun, 20 Sep 2026 14:25:44 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6aaf9e9d355dadd7ed746dca/3215eb31-a24e-4259-a0f0-f0febc05d761.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Intro</h1>
<p>A few years ago on a quest of learning Go, I followed <a href="https://interpreterbook.com/">Writing An Interpreter In Go</a> and built Monkey, a small interpreted programming language. Recently, I came back to it with a different goal: I wanted to see how far I could take the language as an actual development environment.</p>
<p>My plan is to build the tooling around Monkey step by step: first syntax highlighting, then LSP support, and eventually use the language to solve an Advent of Code problem.</p>
<p>As we begin writing code in Monkey, we immediately notice that it lacks the editor features we have come to take for granted: syntax highlighting, diagnostics, navigation, completion, and more. There is a lot of machinery behind features we take for granted, and I wanted to understand it by building it. This post is the first step: teaching editors how to understand Monkey's syntax using Tree-sitter, primarily for highlighting purposes.</p>
<p>Before building the tooling, we should first understand the language we are building it for. Monkey is small, but it has enough language features to build simple tooling around it that would teach us all the basics of parsing. The core functionalities and features in Monkey are the following:</p>
<ul>
<li><p><strong>Data Types:</strong> integers, booleans, strings, arrays, and hash maps.</p>
</li>
<li><p><strong>Variables:</strong> created using <code>let</code> statements (e.g., <code>let foo = 1;</code>).</p>
</li>
<li><p><strong>Operators:</strong> prefix (e.g., <code>-</code>, <code>!</code>) and infix operators (e.g., <code>+</code>, <code>-</code>, <code>*</code>, <code>/</code>, <code>==</code>, <code>!=</code>, <code>&lt;</code>, <code>&gt;</code>) to evaluate arithmetic and boolean expressions.</p>
</li>
<li><p><strong>Functions as First-Class Citizens:</strong> Functions can be bound to names, passed as arguments, and returned from other functions, allowing for higher-order functions.</p>
</li>
<li><p><strong>Closures:</strong> fully supports lexical closures, meaning functions can capture and retain access to the variables in which they were defined.</p>
</li>
<li><p><strong>Built-in Functions:</strong> includes pre-defined utility functions, such as <code>len()</code> (to get the length of strings or arrays), <code>puts()</code> (for printing to standard output), <code>first()</code>, <code>rest()</code>, <code>last()</code>, and <code>push()</code>.</p>
</li>
</ul>
<p>You can find the full docs of Monkey and full Tree-sitter implementation in the <a href="https://github.com/SegniAT/monkey-language-interpreter">Monkey repository</a>.</p>
<p>By the end of this post we'll have a Tree-sitter grammar that parses every construct in Monkey, tested against a corpus, and wired into Neovim.</p>
<hr />
<h1>Why Tree-sitter?</h1>
<p>I decided on using it instead of the approaches most editors use for syntax highlighting before Tree-sitter, such as regular expressions, custom parsers, and lexer-based systems, because Tree-sitter has become widely adopted for editor tooling and syntax-aware features.</p>
<p>There are a few other reasons:</p>
<ul>
<li><p><strong>Error recovery</strong>: Tree-sitter can parse incomplete or syntactically invalid code while still producing a useful syntax tree. This is important for editors because code is often temporarily broken while we are in the process of writing it.</p>
</li>
<li><p><strong>A real syntax tree</strong>: Instead of treating source code as a collection of patterns to match, Tree-sitter produces a syntax tree, which the CLI displays as S-expressions. Editors can use this structural information to provide more accurate, context-aware highlighting. For example, an identifier can be distinguished based on whether it represents a type, variable, function, or something else.</p>
</li>
</ul>
<p>And syntax highlighting is only one of the things we can build on top of a syntax tree. Once we have one, we can also implement features such as:</p>
<ul>
<li><p>Incremental parsing</p>
</li>
<li><p>Code folding</p>
</li>
<li><p>Structural selection</p>
</li>
<li><p>Code navigation</p>
</li>
<li><p>Semantic highlighting</p>
</li>
<li><p>LSP integration and more</p>
</li>
</ul>
<hr />
<h1>Project setup</h1>
<p>Before we start writing the grammar, let's get a minimal Tree-sitter project running. Since the goal of this post isn't to explain the installation process, I'll keep this section short and link to the official setup guide.</p>
<p>Dependencies:</p>
<ul>
<li><p>A JavaScript runtime</p>
</li>
<li><p>A C compiler</p>
</li>
<li><p>Tree-sitter CLI (0.26.9)</p>
</li>
</ul>
<p>Follow the getting started <a href="https://tree-sitter.github.io/tree-sitter/creating-parsers/1-getting-started.html#installation">documentation</a> to set up a similar project.</p>
<hr />
<h1>Building the grammar</h1>
<p>With the Tree-sitter project set up, the next step is to teach it what Monkey looks like. We already have a parser and AST that define the language's structure in the interpreter, so rather than designing the grammar from scratch, we'll translate those existing concepts into Tree-sitter rules.</p>
<h2>1. Starting with a basic grammar</h2>
<h3>The grammar skeleton</h3>
<p>The following is the minimal skeleton of our grammar.</p>
<p><a href="https://github.com/SegniAT/monkey-language-interpreter/blob/main/tree-sitter-monkey/grammar.js"><code>grammar.js</code></a></p>
<pre><code class="language-javascript">/**
 * @file Monkey grammar for tree-sitter
 * @author SegniAT &lt;se.segni.adeba@gmail.com&gt;
 * @license MIT
 */

/// &lt;reference types="tree-sitter-cli/dsl" /&gt;
// @ts-check

export default grammar({
  name: "monkey",

  rules: {
    source_file: $ =&gt; repeat($._statement),
	
    _statement: $ =&gt; choice(
      $.expression_statement
    ),
	
    expression_statement: $ =&gt; seq(
      $.expression,
      optional(';')
    ),
	
	expression: $ =&gt; choice(
      $.identifier,
      $.integer,
	),
	
    identifier: _ =&gt; /[a-zA-Z_]+/,
    integer: _ =&gt; /\d+/,
  }
});
</code></pre>
<p>The grammar function is the most important part of this file. It contains the declarative schema for our language. The <code>name</code> property is the name of the language we're writing the grammar for and the <code>rules</code> property allows us to define rules using built-in functions listed in the <a href="https://tree-sitter.github.io/tree-sitter/creating-parsers/2-the-grammar-dsl.html">documentation</a>.</p>
<p>We build the grammar top-down, mirroring how the interpreter's AST (Abstract Syntax Tree) is organized. One important fact to know up front is that the <strong>start rule</strong> for the grammar is the first property in the rules object. In the example above, that would correspond to <code>source_file</code>, but it can be named anything.</p>
<p>Every grammar rule is written as a JavaScript function that takes a parameter <code>$</code>. The syntax <code>$.identifier</code> is how you refer to another grammar symbol within a rule.</p>
<p>In the <code>monkey</code> parser, a program is a struct that has a slice of statements:</p>
<pre><code class="language-go">// ast/ast.go
type Program struct {
	Statements []Statement
}
</code></pre>
<p>We replicated this in our grammar as follows:</p>
<pre><code class="language-javascript">source_file: $ =&gt; repeat($._statement), // 'repeat(rule)' creates a rule that matches zero-or-more occurrences of a given rule.
</code></pre>
<p>Now we have to define <code>_statement</code>.</p>
<p>In our Monkey parser, there are 3 types of Statements:</p>
<pre><code class="language-go">// ast/ast.go
type Statement interface {
	Node
	statementNode()
}

// 1. The LET statement 
type LetStatement struct {
	Token token.Token // the token.LET token
	Name  *Identifier
	Value Expression
}

// 2. The RETURN statement
type ReturnStatement struct {
	Token       token.Token // the 'return' token
	ReturnValue Expression
}

// 3. The ExpressionStatement statement
type ExpressionStatement struct {
	Token      token.Token // the first token of the expression
	Expression Expression
}
</code></pre>
<p>In our skeleton grammar we only define the Expression Statement for now:</p>
<pre><code class="language-javascript">// Starting a rule's name with an underscore causes the rule to be hidden in the syntax tree. This avoids depth and noise to the syntax tree.
// 'choice(rule1, rule2, ...)' function creates a rule that matches one of a set of possible rules.
_statement: $ =&gt; choice(
      $.expression_statement
	  // We later add the LET and RETURN statements here.
),
</code></pre>
<p>Expression Statement is defined as the following in our grammar:</p>
<pre><code class="language-javascript">// `seq(rule1, rule2, ...)` function creates a rule that matches any number of other rules, in order.
expression_statement: $ =&gt; seq(
	$.expression,
	optional(';') // `optional(rule)` function creates a rule that matches zero or one occurrence of a given rule.
)
</code></pre>
<p>An Expression Statement is just an Expression followed by an optional semicolon (semicolons are optional in Monkey).</p>
<p>Let's look at Expression:</p>
<pre><code class="language-javascript">expression: $ =&gt; choice(
      $.identifier,
      $.integer,
),
	
identifier: _ =&gt; /[a-zA-Z_]+/,
integer: _ =&gt; /\d+/,
</code></pre>
<p>An expression can be an identifier or an integer here, we will add much more to our final version. We describe <code>identifier</code> and <code>integer</code> as regular expressions, an identifier can only include letters and underscores, while an integer only includes numbers.</p>
<p>At this point we have enough grammar to parse something simple, so let's verify that Tree-sitter produces the tree we expect before adding more rules.</p>
<h3>Testing the grammar</h3>
<p>Create <code>test/corpus/basics.txt</code> with these two tests:</p>
<pre><code class="language-Text">=================================
Expression statement (identifier)
=================================

foo

---

(source_file
  (expression_statement
    (identifier)))

==============================
Expression statement (integer)
==============================

42

---

(source_file
  (expression_statement
    (integer)))
</code></pre>
<p>Let's now generate a parser from the grammar and test it against the test.</p>
<pre><code class="language-shell">tree-sitter generate
tree-sitter test
</code></pre>
<p>Output:</p>
<pre><code class="language-plaintext">basics:
    1. ✓ Expression statement (identifier)
    2. ✓ Expression statement (integer)

Total parses: 2; successful parses: 2; failed parses: 0; success percentage: 100.00%; average speed: 593 bytes/ms
</code></pre>
<p>Our generated parser based on the grammar outputs an S-expression when parsing our target language source code, so we use that fact to write our tests. Let's look at the first tests.</p>
<ul>
<li><p>The <strong>name</strong> of each test is written between two lines containing only = (equal sign) characters.</p>
</li>
<li><p>Then the <strong>input source code</strong> is written, followed by a line containing three or more - (dash) characters.</p>
</li>
<li><p>Then, the <strong>expected output syntax tree</strong> is written as an S-expression. The exact whitespace in the S-expression doesn't matter.</p>
</li>
</ul>
<p>As shown in the expected output, the root of our syntax tree is a named node called <code>source_file</code>, which directly corresponds to our root rule defined in <code>grammar.js</code>.</p>
<p>The next node we might expect would be <code>_statement</code>, but the underscore makes it hidden. So we move on to its children, for now we just have <code>expression_statement</code> which has either <code>identifier</code> or <code>integer</code> as its children. In our first test case, <code>foo</code> will be identified as <code>identifier</code>, but in the second <code>42</code> is an <code>integer</code> node.</p>
<h2>2. Building out the grammar</h2>
<p>The basic grammar works, so now we can start filling in the pieces we left out of the skeleton.</p>
<h3>Statements: <code>let</code> and <code>return</code></h3>
<p>Our interpreter has two more Statement types in addition to the already defined Expression Statement. We add them as the other two alternatives of <code>_statement</code> rule:</p>
<pre><code class="language-javascript">// `field(name, rule)` function assigns a field name to the child node(s) matched by the given rule. We can use it to access specific children in the resulting syntax tree.
let_statement: $ =&gt; seq(
  'let',
  field('name', $.identifier),
  '=',
  field('value', $.expression),
  optional(';')
),

return_statement: $ =&gt; seq(
  'return',
  field('value', $.expression), 
  optional(';')
),
</code></pre>
<p>This mirrors the interpreter's AST: the <code>LetStatement</code> struct has a <code>Name</code> and a <code>Value</code> field.</p>
<p>With the remaining Statements in place, we can move on to the remaining Literals and Expressions. This is where the grammar starts getting more interesting. Monkey has several kinds of Literals and Expressions, and some Tree-sitter features become necessary to keep the resulting tree useful and less noisy.</p>
<h3>Literals and Expressions</h3>
<p>Some of the remaining Literals are as follows:</p>
<pre><code class="language-javascript">// `token(rule)` function marks the given rule as producing only a single token. Tree-sitter's default is to treat each `String` or `RegExp` literal in the grammar as a separate token. We don't want 3 separate tokens here, just one.
string: _ =&gt; token(seq('"', /[^"]*/, '"')), // We don't allow '"' in strings since we cannot escape characters in Monkey at the time of writing this.
boolean: _ =&gt; choice("true", "false"),

// ... function, hash, array literals
</code></pre>
<p>In the Monkey grammar, the <code>expression</code> rule is defined as a choice between many different kinds of Expressions:</p>
<pre><code class="language-javascript">    expression: $ =&gt; choice(
      $.identifier,
      $.integer,
      $.string,
      $.boolean,
	  
      $.unary_expression,
      $.binary_expression,
      $.paren_expression,
	  
      $.call_expression,
      $.index_expression,
	  
      $.if_expression,
      $.function_literal,
      $.array_literal,
      $.hash_literal,
),
</code></pre>
<h3>Keyword extraction using <code>word</code></h3>
<p>Consider the following snippet:</p>
<pre><code class="language-Monkey">iffoo
</code></pre>
<p>Tree-sitter would lex this source code as follows:</p>
<ul>
<li><p>an <code>if</code> keyword</p>
</li>
<li><p>an identifier <code>foo</code></p>
</li>
</ul>
<p>But <code>if</code> should only be matched if it appears as a whole word, on its own.</p>
<p>The <code>word</code> property solves this problem. It tells Tree-sitter which rule represents the language's "word" (its identifier). This is what drives <strong>keyword extraction</strong>: Tree-sitter scans the grammar for string literals that could collide with identifiers (<code>let</code>, <code>fn</code>, <code>true</code>, <code>false</code>, <code>if</code>, <code>else</code>, <code>return</code>) and turns them into keywords that <strong>only match as whole words</strong>.</p>
<p>Add <code>word</code> property as a grammar-level setting:</p>
<pre><code class="language-javascript">word: $ =&gt; $.identifier,
</code></pre>
<h3>Abstract categories using <code>supertypes</code></h3>
<p>Some rules in your grammar will represent <strong>abstract categories</strong> of syntax nodes, such as "expression", "type", or "declaration". These rules are often defined as simple choices between several other rules. As shown above, in our grammar, the <code>expression</code> rule is a choice between 13 different rules.</p>
<p>By default Tree-sitter will generate a visible node type for each of these abstract category rules, which can lead to unnecessarily deep and complex syntax trees. To avoid this you can add these abstract category rules to the grammar's <code>supertypes</code> definition. Tree-sitter will then treat these rules as <code>supertypes</code> and will not generate visible node types for them in the syntax tree.</p>
<p><code>supertypes</code> is a cleanup setting, we add it as a grammar-level setting:</p>
<pre><code class="language-javascript">supertypes: $ =&gt; [
  $.expression,
],
</code></pre>
<p><strong>What's the difference between</strong> <code>supertypes</code> <strong>and hidden rules (using</strong> <code>_</code> <strong>prefix)?</strong> Both <code>supertypes</code> and hidden rules are used to keep the final syntax tree cleaner, they are both hidden. However, they serve different purposes and are used in distinct ways.</p>
<p><code>supertypes</code> mark abstract categories as conceptual groups, so they don't appear as concrete nodes while hidden rules hide intermediate helper rules that exist only to structure the grammar, not to represent a meaningful language construct.</p>
<p>A query can target a <code>supertype</code>, and it will automatically match all of its sub-types. When it comes to hidden rules, since the node doesn't exist in the tree, you can't query for it.</p>
<h3>Expression precedence (Pratt precedence)</h3>
<p>This is the most interesting part of our parser, because the interpreter's Pratt parser and Tree-sitter's GLR parser solve the same problem in completely different ways.</p>
<p>Let's look at why we need to explicitly define precedence for rules that may cause conflict.</p>
<p>Let's say our grammar has the following snippet:</p>
<pre><code class="language-javascript">{
  // ...
  expression: $ =&gt; choice(
    $.identifier,
    $.unary_expression,
    $.binary_expression,
    // ...
  ),

  unary_expression: $ =&gt; choice(
    seq('-', $.expression),
    seq('!', $.expression)
  ),

  binary_expression: $ =&gt; choice(
    seq($.expression, '*', $.expression),
    seq($.expression, '+', $.expression),
    // ...
  ),
}
</code></pre>
<p>This flat structure is highly ambiguous. If we try to generate a parser with the <code>tree-sitter generate</code> command Tree-sitter gives us an error message:</p>
<pre><code class="language-Text">Error: Unresolved conflict for symbol sequence:

  '-'  _expression  •  '*'  …

Possible interpretations:

  1:  '-'  (binary_expression  _expression  •  '*'  _expression)
  2:  (unary_expression  '-'  _expression)  •  '*'  …

Possible resolutions:

  1:  Specify a higher precedence in `binary_expression` than in the other rules.
  2:  Specify a higher precedence in `unary_expression` than in the other rules.
  3:  Specify a left or right associativity in `unary_expression`
  4:  Add a conflict for these rules: `binary_expression` `unary_expression`
</code></pre>
<p><code>•</code> in the error message indicates where exactly during parsing the conflict occurs. For an expression like <code>-a * b</code>, it's not clear whether the <code>-</code> operator applies to the <code>a * b</code> or just to the <code>a</code>.</p>
<p>This is where the <code>prec</code> function comes into play. By wrapping a rule with <code>prec</code>, we can indicate that certain sequence of symbols should bind to each other more tightly than others. For example, the <code>-</code>, <code>$.expression</code> sequence in <code>unary_expression</code> should bind more tightly than the <code>$.expression</code>, <code>+</code>, <code>$.expression</code> sequence in <code>binary_expression</code>.</p>
<p>Our interpreter resolves operator precedence <em>in code</em> at runtime, with a table of constants. They are numbered starting from 1 (<code>LOWEST</code>) to 8 (<code>INDEX</code>): <a href="https://github.com/SegniAT/monkey-language-interpreter/blob/main/parser/parser.go#L11"><code>parser/parser.go</code></a></p>
<pre><code class="language-go">const (
	_ int = iota
	LOWEST
	EQUALS      // ==
	LESSGREATER // &gt; or &lt;
	SUM         // +                
	PRODUCT     // *
	PREFIX      // -X or !X
	CALL        // myFunction(X)
	INDEX       // array[index]
)
</code></pre>
<p>Tree-sitter handles precedence <em>in the grammar</em>, by attaching <code>prec</code> and <code>prec.left</code> to the rules themselves. The mapping is almost one-to-one:</p>
<table>
<thead>
<tr>
<th>Operator</th>
<th>Interpreter constant</th>
<th>Tree-sitter</th>
</tr>
</thead>
<tbody><tr>
<td><code>==</code>, <code>!=</code></td>
<td><code>EQUALS</code></td>
<td><code>prec.left(1)</code></td>
</tr>
<tr>
<td><code>&lt;</code>, <code>&gt;</code></td>
<td><code>LESSGREATER</code></td>
<td><code>prec.left(2)</code></td>
</tr>
<tr>
<td><code>+</code>, <code>-</code></td>
<td><code>SUM</code></td>
<td><code>prec.left(3)</code></td>
</tr>
<tr>
<td><code>*</code>, <code>/</code></td>
<td><code>PRODUCT</code></td>
<td><code>prec.left(4)</code></td>
</tr>
<tr>
<td><code>-x</code>, <code>!x</code> (prefix)</td>
<td><code>PREFIX</code></td>
<td><code>prec(5)</code></td>
</tr>
<tr>
<td><code>f(x)</code> (call)</td>
<td><code>CALL</code></td>
<td><code>prec.left(7)</code></td>
</tr>
<tr>
<td><code>a[i]</code> (index)</td>
<td><code>INDEX</code></td>
<td><code>prec.left(8)</code></td>
</tr>
</tbody></table>
<pre><code class="language-javascript">{
  // ...

// `prec(number, rule)` function marks the given rule with a numerical precedence, which will be used to resolve LR(1) Conflicts (https://en.wikipedia.org/wiki/LR_parser#Conflicts_in_the_constructed_tables) at parser-generation time.
  unary_expression: $ =&gt;
    prec(
      5,
      choice(
        seq("-", $.expression),
        seq("!", $.expression),
        // ...
      ),
    );

	// `prec.left([number], rule)` function marks the given rule as left-associative (and optionally applies a numerical precedence).
	binary_expression: $ =&gt; choice(
      prec.left(1, seq(field('left',$.expression), '==', field('right',$.expression))),
      prec.left(1, seq(field('left',$.expression), '!=', field('right',$.expression))),
      prec.left(2, seq(field('left',$.expression), '&lt;', field('right',$.expression))),
      prec.left(2, seq(field('left',$.expression), '&gt;', field('right',$.expression))),
      prec.left(3, seq(field('left',$.expression), '+', field('right',$.expression))),
      prec.left(3, seq(field('left',$.expression), '-', field('right',$.expression))),
      prec.left(4, seq(field('left',$.expression), '*', field('right',$.expression))),
      prec.left(4, seq(field('left',$.expression), '/', field('right',$.expression))),
    ),
}
</code></pre>
<p><code>prec.left(n)</code> means: this operator binds with precedence <code>n</code> and is <strong>left-associative</strong>, so <code>a - b - c</code> parses as <code>(a - b) - c</code>, exactly like the interpreter's loop, which keeps parsing while the next token's precedence is higher. Prefix operators use plain <code>prec(n)</code>, because associativity doesn't apply to them.</p>
<p>In the Pratt parser the constants only need to be <em>relatively</em> ordered. Tree-sitter's numbers do the same job, so the two tables line up almost exactly.</p>
<p>There is one <strong>main difference</strong> between the two styles worth knowing: the Pratt parser decides how to continue looking forward from an expression, but the GLR parser explores <em>all</em> parse branches in parallel and only uses precedence to resolve conflicts when they actually arise. That's why for example <code>choice</code> order doesn't matter in Tree-sitter.</p>
<h3>Calls, indexing, and the remaining constructs</h3>
<p>With operator precedence sorted out, we can finish the remaining Expressions. Calls and indexing are Infix Expressions with the highest binding as seen in the previous section.</p>
<pre><code class="language-javascript">call_expression: $ =&gt; prec.left(7, seq(
  field('function', $.expression),
  field('arguments', $.argument_list)
)),

index_expression: $ =&gt; prec.left(8, seq(
  field('object', $.expression),
  '[',
  field('index', $.expression),
  ']',
)),
</code></pre>
<p>The remaining constructs need no new concepts, they are combinations of the primitives we have already seen.</p>
<pre><code class="language-javascript">block: $ =&gt; seq('{', repeat($._statement), '}'),

if_expression: $ =&gt; seq(
  'if',
  '(',
  field('condition', $.expression),
  ')',
  field('consequence', $.block),
  optional(seq('else', field('alternative', $.block))),
),

function_literal: $ =&gt; seq(
  'fn',
  field('parameters', $.parameter_list),
  field('body', $.block)
),

hash_pair: $ =&gt; seq(
  field('key', $.expression),
  ':',
  field('value', $.expression),
)
</code></pre>
<h2>3. The complete grammar</h2>
<p>At this point we've covered all the Tree-sitter concepts that required explanation. The remaining rules are mostly combinations of the same primitives, so we can jump directly to the complete grammar. The final <code>grammar.js</code> is only <a href="https://github.com/SegniAT/monkey-language-interpreter/blob/main/tree-sitter-monkey/grammar.js">~145 lines</a>.</p>
<p>Feeding it a complete program shows what we built. Create a file with the following content:</p>
<pre><code class="language-Text">let x = 5;
let add = fn(a, b) { return a + b; };
let result = add(3, 4);
let arr = [1, 2, 3, 4];
let person = {"name": "Alice", "age": 30};
let age = person["age"];
if (age &gt; 21) { let status = "adult"; } else { let status = "minor"; }
let makeAdder = fn(x) { return fn(y) { return x + y; }; };
let addFive = makeAdder(5);
addFive(10);
</code></pre>
<p>Running the command <code>tree-sitter parse [your file name]</code> should output the following:</p>
<pre><code class="language-Text">(source_file
  (let_statement name: (identifier) value: (integer))
  (let_statement
    name: (identifier)
    value: (function_literal
      parameters: (parameter_list
        name: (identifier)
        name: (identifier))
      body: (block
        (return_statement
          value: (binary_expression
            left: (identifier)
            right: (identifier))))))
  (let_statement
    name: (identifier)
    value: (call_expression
      function: (identifier)
      arguments: (argument_list
        (integer)
        (integer))))
  (let_statement
    name: (identifier)
    value: (array_literal
      (integer) (integer) (integer) (integer)))
  (let_statement
    name: (identifier)
    value: (hash_literal
      (hash_pair
        key: (string)
        value: (string))
      (hash_pair
        key: (string)
        value: (integer))))
  (let_statement
    name: (identifier)
    value: (index_expression
      object: (identifier)
      index: (string)))
  (expression_statement
    (if_expression
      condition: (binary_expression
        left: (identifier)
        right: (integer))
      consequence: (block
        (let_statement
          name: (identifier)
          value: (string)))
      alternative: (block
        (let_statement
          name: (identifier)
          value: (string)))))
  (let_statement
    name: (identifier)
    value: (call_expression
      function: (identifier)
      arguments: (argument_list
        (integer))))
  (expression_statement
    (call_expression
      function: (identifier)
      arguments: (argument_list
        (integer)))))
</code></pre>
<p>Every construct shows up as a named node, with fields (<code>name:</code>, <code>value:</code>, <code>left:</code>, <code>right:</code>, <code>condition:</code>, <code>consequence:</code>, ...) carrying the same information the interpreter's AST carries in Go.</p>
<p>The grammar is now complete, but that does not mean that the generated parser outputs the correct syntax tree. Validation is our next undertaking.</p>
<hr />
<h1>The corpus as executable specification</h1>
<p><code>tree-sitter generate</code> compiles the grammar into a C parser, and <code>tree-sitter test</code> runs the corpus (test cases). There are 70 test cases in the <a href="https://github.com/SegniAT/monkey-language-interpreter/tree/main/tree-sitter-monkey/test/corpus">project</a> written alongside the grammar:</p>
<pre><code class="language-plaintext">Total parses: 70; successful parses: 70; failed parses: 0; success percentage: 100.00%
</code></pre>
<p><strong>The corpus is the spec.</strong> Because each test includes an input source code and the expected output tree, it also serves as documentation of the language.</p>
<p><strong>Error recovery</strong>: Tree-sitter's superpower for editors is its behavior on broken code. Instead of stopping the parsing task, the parser emits an <code>ERROR</code> node and continues its parsing. The tree stays mostly correct while you type. We test these cases in the corpus with the <code>:error</code> directive:</p>
<pre><code class="language-Text">===========================
Integer as let name
:error
===========================

let 5 = x;

---

============================
Missing value
:error
============================

let x = ;

---

============================
Unterminated list
:error
============================

[1, 2
</code></pre>
<p>These tests assert that the input contains a syntax error while still allowing Tree-sitter to produce a useful tree around it. This is why highlighting stays useful in an editor. Code is broken most of the time while we're writing it, and the parser still produces a useful tree around the errors.</p>
<p>At this point we have a parser that can turn Monkey source code into a syntax tree, and we can test that parser against a corpus to verify its validity. But our original goal wasn't just to parse Monkey, we wanted an editor to understand it well enough to highlight it.</p>
<hr />
<h1>Turning the parsed tree into highlighting</h1>
<p>To achieve our main goal of syntax highlighting, we need to write queries that match nodes in the syntax tree and assign them <em>captures</em>, those are names that Neovim understands and maps to highlight groups it has already defined.</p>
<p>You can find the list of standard captures supported by Neovim in their <a href="https://neovim.io/doc/user/treesitter/#treesitter-highlight-groups">Tree-sitter docs</a>. These include captures such as <code>@function</code>, <code>@variable</code>, <code>@keyword</code>, and <code>@string</code>, which are highlighted according to the current colorscheme.</p>
<p>The queries assigned to captures live at <a href="https://github.com/SegniAT/monkey-language-interpreter/blob/main/tree-sitter-monkey/queries/highlights.scm"><code>queries/highlights.scm</code></a>. The ordering inside it matters, more specific patterns need to come before generic ones. But we will set priorities using <code>#set! priority N</code> to be more explicit. Higher numbers win, which lets us keep related patterns grouped together instead of carefully ordering every rule.</p>
<p>In our case, we have different types of functions and identifiers:</p>
<pre><code class="language-Scheme">; User-defined function calls
((call_expression
  function: (identifier) @function.call)
  (#set! priority 120))

; Built-in function calls
((call_expression
  function: (identifier) @function.builtin
  (#any-of? @function.builtin "len" "first" "last" "rest" "push" "puts"))
  (#set! priority 130))

; Function parameters
((function_literal
  parameters: (parameter_list
    (identifier) @variable.parameter))
  (#set! priority 120))

; Hash keys are properties (strings and bare identifiers)
((hash_pair
  key: (string) @property)
  (#set! priority 120))

((hash_pair
  key: (identifier) @property)
  (#set! priority 120))

; All-caps identifiers are constants by convention, which is a convention from other languages - Monkey has no reassignment, so every variable is effectively constant.
((identifier) @constant
  (#match? @constant "^[A-Z][A-Z_]*$")
  (#set! priority 120))

; more captures ...

; Catch-all: every other identifier is a variable
(identifier) @variable
</code></pre>
<p>Two <a href="https://tree-sitter.github.io/tree-sitter/using-parsers/queries/3-predicates-and-directives.html">predicates</a> do the heavy lifting here:</p>
<ul>
<li><p><code>#any-of?</code> checks the node's text against a list, that's how <code>len</code> and other builtin functions become <code>@function.builtin</code> while every other callee is <code>@function.call</code>. Monkey has no named function definitions (functions assigned to a variable with <code>let</code>), so <code>@function</code> itself stays unused.</p>
</li>
<li><p><code>#match?</code> applies a regex, all-caps identifiers become <code>@constant</code> by convention, before the catch-all below turns everything else into <code>@variable</code>. The Monkey language does not have the concept of constant variables, this is just for convention.</p>
</li>
</ul>
<p>The rest is a straightforward mapping of literals and keywords:</p>
<pre><code class="language-Scheme">; Keywords
"let" @keyword

[
  "if"
  "else"
] @keyword.conditional

"return" @keyword.return

"fn" @keyword.function

; Literals
[
  "true"
  "false"
] @boolean

(integer) @number

(string) @string

; ... more
</code></pre>
<p>Keywords and literals map directly, and the rest of <a href="https://github.com/SegniAT/monkey-language-interpreter/blob/main/tree-sitter-monkey/queries/highlights.scm"><code>queries/highlights.scm</code></a> covers operators, brackets, and delimiters (<code>,</code>, <code>;</code>, <code>:</code>). Check out the file to explore all the captures. The full list is 18 captures which is tiny compared to languages like TypeScript or Rust for example.</p>
<p>We now have both pieces we need to finalize our project: a parser that produces the tree and highlight queries that turn nodes into captures. The final step is connecting that grammar to Neovim so it knows when to use the parser for .monkey files.</p>
<hr />
<h1>Neovim integration and showcase</h1>
<p>Now comes our final task, i.e. editor integration. We will follow the <a href="https://github.com/nvim-treesitter/nvim-treesitter/blob/main/README.md">nvim-treesitter README.md</a>, "Adding custom languages" section.</p>
<p>First, we need to let Neovim know that <code>.monkey</code> files are a thing, in our configuration file we add the following:</p>
<pre><code class="language-Lua">vim.filetype.add({ extension = { monkey = "monkey" } })
</code></pre>
<p>Then we need to register the parser with <code>nvim-treesitter</code>:</p>
<pre><code class="language-Lua">vim.api.nvim_create_autocmd('User', {
  pattern = 'TSUpdate',
  callback = function()
    require('nvim-treesitter.parsers').monkey = {
      install_info = {
        url = 'https://github.com/SegniAT/monkey-language-interpreter',
        location = 'tree-sitter-monkey',
        queries = 'tree-sitter-monkey/queries',
      },
    }
  end,
})
</code></pre>
<p>The interesting part is the monorepo layout: our grammar lives <em>inside</em> the interpreter's repository. nvim-treesitter downloads the repo tarball from <code>url</code>, builds the grammar found at <code>location</code>, and installs the queries found at <code>queries</code>.</p>
<p>This works, but it isn't really ideal since the grammar is only one directory inside a much larger repository of the Monkey interpreter (and soon, the LSP as well)! So installing the parser means downloading the entire interpreter repository.</p>
<p>After this, the update loop is:</p>
<pre><code class="language-shell">tree-sitter generate &amp;&amp; tree-sitter test   # locally
git commit &amp;&amp; git push                     # to the repo
:TSUpdate monkey                           # in Neovim
</code></pre>
<p>That's enough to get the whole pipeline working. We can now open a .monkey file in Neovim, have Tree-sitter parse it, apply our queries, and get syntax highlighting.</p>
<p><strong>Before:</strong></p>
<img src="https://cdn.hashnode.com/uploads/covers/6aaf9e9d355dadd7ed746dca/c4941b41-dcae-4eaf-a671-93710af11b6a.png" alt="Monkey code before highlighting." style="display:block;margin:0 auto" />

<p><strong>After:</strong> (<code>tokyonight-night</code> theme)</p>
<img src="https://cdn.hashnode.com/uploads/covers/6aaf9e9d355dadd7ed746dca/5c875eb9-a27d-4226-808d-7641afd2014e.png" alt="Monkey code with highlighting." style="display:block;margin:0 auto" />

<p><em>Beautiful!</em></p>
<p>A look at the generated tree by the parser using <code>:InspectTree</code> in Neovim:</p>
<p><a class="embed-card" href="https://youtu.be/ezHMVNos348">https://youtu.be/ezHMVNos348</a></p>

<p>A few queries on the generated tree using <code>:EditQuery</code> in Neovim:</p>
<p><a class="embed-card" href="https://youtu.be/ycQNCGB2c9g">https://youtu.be/ycQNCGB2c9g</a></p>

<hr />
<h1>Conclusion</h1>
<p>To add syntax highlighting and make our custom language prettier to look at and more convenient to work with, we had to go through a whole bunch of fun challenges: a hand-written grammar, a corpus that also serves as a spec, queries that map the tree to editor semantics, and finally editor integration.</p>
<p>The biggest lesson: <strong>the grammar is an executable spec and the corpus is its test suite</strong>. Together they make the language's actual behavior much harder to misunderstand than documentation alone.</p>
<p>A second lesson: <strong>a toy language is still a precise language.</strong> Monkey's grammar fits in ~145 lines, yet every rule had to be right: precedence levels, hidden rules, keyword extraction, fields, this is the same machinery that parses Go or Zig, just much smaller.</p>
<p>But syntax highlighting is only the beginning. We can now tell the editor what a piece of code looks like, but not much about what it means. The next step is to teach the editor about Monkey itself: diagnostics, definitions, completion, and eventually the other features we expect from a modern development environment. That's where the LSP comes in, in our second entry the series.</p>
<hr />
<h1>References</h1>
<ul>
<li><p><a href="https://interpreterbook.com/">Writing An Interpreter In Go</a> by Thorsten Ball</p>
</li>
<li><p><a href="https://tree-sitter.github.io/tree-sitter/">Tree-sitter documentation</a></p>
</li>
<li><p><a href="https://github.com/SegniAT/monkey-language-interpreter">The Monkey interpreter + tree-sitter grammar repository</a></p>
</li>
<li><p><a href="https://youtu.be/09-9LltqWLY">TJ DeVries - Tree-sitter explained</a></p>
</li>
<li><p><a href="https://youtu.be/a1rC79DHpmY">Tree-sitter: a new parsing system for programming tools - GitHub Universe 2017</a></p>
</li>
</ul>
]]></content:encoded></item></channel></rss>