Files
mist-website/docs.md
T
2026-07-03 10:14:43 +02:00

27 KiB

Mist Language Documentation

Introduction

Mist is a pragmatic systems programming language that compiles directly to Rust. It provides C/C++-style syntax with zero-cost abstractions, full interoperability with the Rust ecosystem, and additional language features like classes with inheritance.

The compiler is written in Rust and Mist itself, and is organized into four main crates:

  • mist-parser — Parses Mist source code into an AST using pest
  • mist-codegen — Generates Rust code from the Mist AST
  • mist-analyzer — Language server built on top of rust-analyzer
  • mist-api — Mist crate that orchestrates transpilation and building

Installation

cargo install mist-lang

This installs the mist CLI binary and the mist-analyzer LSP binary.


Project Structure

A Mist project requires the following structure:

my-project/
├── Mist.toml          # Project configuration
├── Cargo.toml         # Standard Rust Cargo.toml (managed by mist)
├── src/
│   ├── main.mist      # Entry point (or lib.mist for libraries)
│   └── ...            # Other .mist source files
└── .mist/
    └── src/           # Generated Rust output (do not edit)

Mist.toml

package = "main.mist"       # Entry point file (relative to src/)
packages = ["utils"]        # Subdirectories to treat as sub-modules
include = ["crates/*"]      # Additional Mist crates to transpile

The packages field lists subdirectories under src/ that should be treated as child modules. Each such directory should contain a package.mist file acting as the module root.

The include field supports glob patterns for including external Mist crates. A trailing * treats all immediate subdirectories that contain a Mist.toml as independent crates.


CLI Commands

Command Short Description
mist run r Run the project
mist build b Build the project
mist check c Check without building
mist test Run tests
mist publish Publish the project
mist bench Benchmark the project
mist doc Build documentation
mist fix Auto-fix warnings
mist clippy Lint the project
mist clean Clean build artifacts
mist transpile t Transpile Mist to Rust only
mist init Init a project in current dir
mist new Create a new project
mist version -v Print compiler version
mist help -h Print usage information

All Cargo subcommands delegate to cargo with Mist error remapping.


Language Syntax

Comments

// Line comments only

Literals

42              // Integer
3.14            // Float
true            // Boolean
false           // Boolean
"hello"         // String
(1, true, "x")  // Tuple

Identifiers & Keywords

Keywords are reserved and cannot be used as identifiers:

if, else, fn, for, while, match, return, break, continue, struct, enum, class, trait, impl, use, pub, mut, let, true, false, dyn, loop, unsafe, override, const, type

Identifiers follow the pattern [a-zA-Z_][a-zA-Z0-9_]*.


Variables

let x = 42;                    // Type-inferred immutable
let mut y = 10;                // Mutable variable
i32 z = 100;                   // Explicit type annotation
str& s = "hello";              // Typed string reference
bool b = true;                 // Typed boolean
f64 f = 3.14;                  // Typed float
let (a, b) = (1, "two");      // Destructuring
let (x, (y, z)) = (1, (2, 3)); // Nested destructuring

Variable declarations follow either of two forms:

let <pattern> [= <expr>];
<type> <pattern> [= <expr>];

Functions

// Return type is `void` (unit)
void greet(str& name)
{
    println!("Hello, {name}");
}

// With return type
i32 add(i32 a, i32 b)
{
    return a + b;
}

// Expression body (last expression is the return value)
i32 square(i32 x)
{
    x * x
}

// Public function
pub i32 multiply(i32 a, i32 b)
{
    a * b
}

// Generic function
T identity<T>(T value)
{
    value
}

// Unsafe function
unsafe i32 dangerous()
{
    42
}

Function syntax: [pub] [void | <type>] <name>[<generics>](<params>) [override] { <body> }

Functions use Allman brace style (opening brace on the next line).

Self Parameter

Methods can take self, &self, &mut self, or self with lifetimes:

pub void set_value(&mut self, i32 v)
{
    self.value = v;
}

Control Flow

If/Else

if condition {
    // body
} else if other_condition {
    // body
} else {
    // body
}

Conditions must be parenthesized: if (expr) { } or if expr_no_struct { } (struct literals require parentheses).

Used as an expression:

let x = if true { 1 } else { 2 };

While

while condition {
    // body
}

For

for pattern in iterator {
    // body
}

for i in 0 .. 10 {
    // body
}

C-Style For

for (let mut i = 0; i < 10; i++) {
    // body
}

Loop

loop {
    // infinite loop
}

Match

match value {
    pattern1 => expr,
    pattern2 => {
        // block body
    }
    pattern3 | pattern4 => expr,
}

Break/Continue

break;
continue;

Return

return;
return value;

Types

Type System

i32                // Path type
str&               // Reference type
&mut i32           // Mutable reference
&'a i32            // Reference with lifetime
*const i32         // Const pointer
*mut i32           // Mutable pointer
(i32, bool)        // Tuple type
fn(i32) -> bool    // Function pointer type
Fn(i32) -> bool    // Closure trait (Fn)
FnMut(i32) -> bool // Closure trait (FnMut)
FnOnce(i32) -> bool // Closure trait (FnOnce)
dyn Trait          // Trait object
void               // Unit type (maps to Rust's ())
'i32               // Type ascription not to be confused with lifetime

Type expressions can be composed:

type_expr = { (void | path_type | tuple_type | dyn_type) ~ (unsafe_ref_type | ref_type | fn_type)* }

This means types are written left-to-right naturally:

i32&          // &i32
i32&mut       // &mut i32
i32&'a        // &'a i32
i32 fn() -> bool  // fn(i32) -> bool

Type Aliases

type MyInt = i32;

type Result<T> = std::result::Result<T, str&>;

Const and Static

const MAX: i32 = 100;
static NAME: str& = "hello";

Structs

pub struct Point {
    i32 x,
    i32 y,
}

pub struct Generic<T> {
    T value,
}

Fields can be public or private:

struct User {
    pub str& name,
    i32 age,  // private
}

Enums

pub enum Option<T> {
    Some(T),
    None,
}

pub enum Message {
    Quit,
    Move(i32, i32),
    Write { str& content, i32 length },
}

Enum variants can be:

  • Named — Variant
  • Tuple — Variant(T1, T2)
  • Struct — Variant { T1 field1, T2 field2 }

Traits

pub trait Drawable {
    void draw(&self);
}

pub trait Comparable<T> : Eq {
    i32 cmp(&self, T other);
}

Trait requirements are specified after ::

trait MyTrait : SuperTrait + OtherTrait {
    void required_method(&self);
    void another(&self);
}

Impl Blocks

impl MyType {
    void method(&self) { }
}

impl Trait for MyType {
    void method(&self) { }
}

impl<T> GenericTrait<T> for MyType {
    void method(&self, T value) { }
}

Classes

Mist introduces class as syntactic sugar for a Rust struct with a virtual method table (vtable).

pub class Animal {
    str& name,

    pub constructor(str& name)
    {
        self.name = name;
    }

    pub void speak(&self)
    {
        println!("...");
    }
}

Class fields can have default initializers:

class Player {
    i32 health = 100,
    str& name,
}

Inheritance

class Dog : Animal {
    pub constructor(str& name)
    {
        super(name);
    }

    override void speak(&self)
    {
        println!("Woof!");
    }
}

The override keyword supports explicit base class targeting:

override(Animal) void speak(&self)
{
    println!("Woof!");
}

Under the Hood

A class Dog : Animal generates:

  1. A Rust struct with a _super: Animal field (or _vptr: &'static [*const c_void] for root classes)
  2. A vtable constant with function pointers for each public method
  3. An impl block with Deref<Target = Animal> and DerefMut
  4. Method trampolines (__m_<name>) that are dispatched through the vtable
  5. A new() constructor that initializes via MaybeUninit and calls the user's constructor(&mut self)

The vtable is unified: parent entries are copied, overridden entries replace parent slots, and new methods are appended.


Generics

// Generic function
T id<T>(T x)
{
    x
}

// Generic struct
struct Pair<A, B> {
    A first,
    B second,
}

// Generic enum
enum Result<T, E> {
    Ok(T),
    Err(E),
}

// Generic with trait bounds
T max<T : Ord>(T a, T b)
{
    if a > b { a } else { b }
}

// Lifetime generics
void process<'a>(&'a str& data)
{
    // ...
}

Generic syntax uses < > delimiters. Lifetimes are prefixed with '.


Visibility

// Private (default)
void internal() { }

// Public
pub void external() { }

// Public to specific path
pub(crate) void crate_only() { }
pub(super) void parent_only() { }
pub(in my::module) void module_only() { }

Modules

// Declare a submodule
pub module foo;

// Import
use std::collections::HashMap;

// Re-exporting import
pub use my_module::MyType;

Module Resolution

Mist maps the module tree to Rust's module system:

Mist Path Rust Output
src/main.mist .mist/src/main.rs
src/utils/package.mist .mist/src/utils/mod.rs
src/foo.mist .mist/src/foo.rs
src/utils/helper.mist .mist/src/utils/helper.rs

A package.mist file acts as a directory's module root, analogous to mod.rs.


Attributes

Inner attributes apply to the containing module:

#![allow(unused_variables)]

Outer attributes apply to the next item:

#[derive(Debug, Clone)]
struct Point {
    i32 x,
    i32 y,
}

#[test]
void my_test()
{
    assert_eq!(1, 1);
}

Attribute syntax:

  • #[path]
  • #[path = literal]
  • #[path(item1, item2, ...)]

Patterns

// Literal patterns
match x {
    1 => "one",
    2 => "two",
    _ => "other",
}

// Tuple patterns
let (a, b) = (1, 2);

// Struct patterns
match value {
    Point { x, y } => x + y,
    Point { x: 0, y } => y,
    _ => 0,
}

// Named tuple patterns (newtype)
let MyType(value) = my_var;

// Wildcard / etc
let _ = get_side_effect();
match x {
    1 => ...,
    .. => ...,  // rest / etc
}

// Mutable binding in pattern
let mut x = 42;
match ref_to_option {
    Some(mut value) => value += 1,
    None => {},
}

Operators

Binary Operators

Operator Description
+ Addition
- Subtraction
* Multiplication
/ Division
% Modulus
== Equality
!= Inequality
< Less than
> Greater than
<= Less or equal
>= Greater or equal
&& Logical AND
|| Logical OR
& Bitwise AND
| Bitwise OR
^ Bitwise XOR
<< Left shift
>> Right shift
= Assignment
+= Add assign
-= Subtract assign
*= Multiply assign
/= Divide assign
%= Modulus assign
&= Bitwise AND assign
|= Bitwise OR assign
^= Bitwise XOR assign
<<= Left shift assign
>>= Right shift assign
.. Range (exclusive)
..= Range (inclusive)
-> Pointer write

Prefix Operators

Operator Description
* Dereference
& Reference
&mut Mutable reference
! Logical NOT
- Numeric negation

Postfix Operators

Operator Description
.field Field access
.0 Tuple field access
() Function call
[] Index
{ f: v } Struct literal
as Type Type cast
? Try (error prop)
++ Increment
-- Decrement
!() Macro call (paren)
![] Macro call (bracket)
!{} Macro call (brace)

Macros

Mist reuses Rust's macro system directly:

println!("hello");
assert_eq!(a, b);
vec![1, 2, 3];

Macro calls use ! followed by parentheses, brackets, or braces.


Closures

let add = (a, b) => a + b;
let result = add(2, 3);  // 5

let square = f64 x => x * x;

Closures can have explicit return types:

let transform = i32 x => {
    x * 2
};

Attributes

Mist supports Rust-style attributes at module and item level:

#![crate_type = "lib"]

#[derive(Clone)]
#[repr(C)]
pub struct External { }

Compiler Architecture

Pipeline Overview

The Mist compiler operates in distinct phases:

Source (.mist)
    │
    ▼
┌─────────────┐
│   Lexer     │  (pest PEG grammar)
│   & Parser  │
└─────────────┘
    │  AST
    ▼
┌─────────────┐
│  Semantics  │  (field init checking)
└─────────────┘
    │
    ▼
┌─────────────┐
│  Codegen    │  (AST → Rust source)
└─────────────┘
    │  .rs + .map.json
    ▼
┌─────────────┐
│    cargo    │  (Rust compilation)
└─────────────┘
    │
    ▼
  Binary / Library

1. Parsing (mist-parser)

The parser is built with pest, a PEG parser generator. Grammar rules are defined in grammar.pest (693 rules).

Key files:

  • grammar.pest — PEG grammar defining the entire language syntax
  • src/parser/common/ — Parse rule → AST conversions for expressions, statements, types, declarations
  • src/parser/items/ — Parse rule → AST conversions for top-level items (structs, enums, classes, functions, traits, impls, attributes)
  • src/ast/ — AST node types (expr.rs, statement.rs, top_level.rs)
  • src/semantics.rs — Semantic checks (e.g., ensuring all class fields are initialized in constructors)
  • src/error.rs — Error types (PreAst parsing errors, Ast generation errors)
  • src/rev_mapper.rs — Position mapping between Mist and Rust source

Parsing entry points:

// Parse a complete program
pub fn parse<'a>(source: &'a str) -> Result<Program, ParseError<'a>>

// Parse only the module declaration
pub fn parse_module<'a>(source: &'a str) -> Result<Option<(Visibility, Identifier)>, ParseError<'a>>

Grammar Structure

The grammar follows a layered approach:

  1. Lexical rules (silent) — WHITESPACE, COMMENT, identifier, integer, float, string_lit
  2. Primary expressions — literals, paths, tuples, arrays, closures, control flow
  3. Term — prefix operators + primary + postfix operators
  4. Expression — terms joined by binary operators (via Pratt parser)
  5. Statements — variable declarations, control flow (if, while, for, match, loop), blocks
  6. Top-level items — functions, structs, enums, classes, traits, impls, type aliases, imports, module declarations, constants

AST Structure

The AST preserves source positions via Spanned<T>:

pub struct Spanned<T> {
    pub line: usize,
    pub column: usize,
    pub item: T,
}

Top-level items are wrapped as:

pub struct TopLevel(pub Spanned<TopLevelKind>, pub Vec<Attribute>);

Expressions use a fix-point representation for prefix/postfix operators:

Expression::Fix {
    initial: Box<Expression>,
    prefixes: Vec<Prefix>,
    postfixes: Vec<Postfix>,
}

Binary expressions use Pratt parsing for correct precedence:

Expression::Binary {
    lhs: Box<Expression>,
    op: String,
    rhs: Box<Expression>,
}

All operators are left-associative with a single precedence level.

2. Semantic Analysis (mist-parser)

The semantic checker (semantics.rs) performs class field initialization analysis. When a class has a constructor, it verifies that every declared field is mutated (directly or indirectly via method calls) within the constructor body.

Key check: check_class_semantics() — Collects all field names, then walks the constructor body tracking which identifiers are assigned. If any field is uninitialized, an error is reported.

The mutability analysis (GetMutability trait) works by:

  1. Collecting all identifiers that get &mut references
  2. Tracking self.field = value patterns
  3. Following method calls that might initialize fields
  4. Ensuring all branches initialize the same fields (intersection semantics)

3. Code Generation (mist-codegen)

The code generator converts the Mist AST into Rust source code. It does not generate an intermediate representation — it produces Rust source text directly.

Key files:

  • src/lib.rs — RustCodegen struct with output buffer, indentation tracking, Mapping for position translation
  • src/top_level.rs — Generates Rust for top-level items (structs, enums, traits, functions, impls, imports, type aliases, const/static)
  • src/statement.rs — Generates Rust for statements and blocks
  • src/expr.rs — Generates Rust for expressions (literals, paths, binary ops, closures, arrays, prefix/postfix)
  • src/class_decl.rs — Class-specific code generation (struct + vtable + impl blocks + Deref)

Codegen Design

The RustCodegen maintains:

  • A string output buffer
  • An indent level (4 spaces per level)
  • A Mapping that records (RustMap, MistMap) pairs for source position remapping
  • A current position tracker (RustMap(line, column))

Generation follows the GenRust trait:

pub trait GenRust {
    fn gen_rust(&self, ctx: &mut Context, cg: &mut RustCodegen);
}

And GetRust for simple string-returning types:

pub trait GetRust {
    fn get_rust(&self) -> String;
}

The Context carries optional expression path information for class super / Super resolution.

Class Codegen (detailed)

Classes are the most complex codegen path. ClassProcessedData analyzes a class declaration and then emits:

  1. Struct declaration — struct ClassName<G> { pub _super: Parent, pub field1: T1, ... } (or _vptr for root classes)

  2. Vtable constants — Index constants __FN_METHOD and a __V_TABLE static array of function pointers. For inherited classes, parent vtable entries are copied and overridden entries replaced.

  3. Constructor — pub fn new(...) -> Self that:

    • Creates an uninitialized instance via MaybeUninit::zeroed().assume_init()
    • Sets the vtable pointer
    • Writes field defaults
    • Calls self.constructor(...)
    • Returns this
  4. Method trampolines — Public methods with self get wrapper functions __m_method that are stored in the vtable, plus virtual dispatch methods that look up the function pointer at runtime.

  5. Override support — Methods marked override are validated at compile time by generating test code that dereferences &Self to &Target, confirming the Deref chain works.

  6. Deref impls — impl Deref<Target = Parent> for Child and DerefMut for inherited classes.

  7. Impl declarations — Inner impl blocks are rewritten to use the self type.

Mapping System

The rev_mapper module provides bidirectional position mapping:

pub struct Mapping {
    pub mist_path: PathBuf,
    pub map: HashSet<(RustMap, MistMap)>,
}
  • RustMap(usize, usize) — line/column in generated Rust
  • MistMap(usize, usize) — line/column in original Mist

The mapping is populated during codegen via GenSpanTranslation:

impl<T> GenSpanTranslation for Spanned<T> {
    fn gen_span(&self, cg: &mut RustCodegen) {
        cg.mapping.map.insert((cg.position, MistMap(self.line, self.column)));
    }
}

This allows the builder to remap Rust compiler errors back to the original Mist source positions.

4. Builder & Error Remapping (mist-api)

The builder module (builder.mist in the mist_api crate) wraps cargo to:

  1. Spawn cargo with --message-format=json
  2. Parse JSON compiler messages
  3. For each diagnostic span, look up the corresponding .map.json file
  4. Remap Rust line/column to Mist line/column using the mapping
  5. Display errors/warnings with Mist source locations and context lines
pub fn build(args: Vec<String>, root: PathBuf) -> bool {
    // Spawn cargo with JSON output
    // Parse CompilerMessage for each span
    // Look up rev_mapper::Mapping from .map.json
    // Remap to Mist positions
    // Print diagnostics
}

5. Transpilation Pipeline

The transpiler (transpiler.mist in mist_api) orchestrates the full pipeline:

  1. Read Mist.toml configuration
  2. Build a Module tree from the filesystem (discovering package.mist files)
  3. For each module: a. Parse Mist source b. Run semantic checks c. Generate Rust code and position mapping d. Write .rs file and .map.json file
  4. Recursively transpile included crates

The Module struct represents the file-to-module mapping:

pub class Module {
    pub String name;
    pub PathBuf path;
    pub Vec<Module> children;

    pub constructor(PathBuf mist_path, MistConfig& config) { ... }
    pub bool is_package(&self) { ... }
    pub PathBuf output_dir(&self, PathBuf& parent_dir) { ... }
    pub PathBuf output_path(&self, PathBuf& parent_dir, ...) { ... }
}

Transpilation Mapping Rules

Mist Source Rust Output
src/main.mist .mist/src/main.rs
src/foo.mist .mist/src/foo.rs
src/bar/package.mist .mist/src/bar/mod.rs
src/bar/baz.mist .mist/src/bar/baz.rs

A package.mist file:

  • Generates mod.rs in the output directory
  • Contains pub mod <child>; declarations for its children
  • Its own items are prepended to the output

6. Caching

The transpiler implements basic caching via file modification times:

fn is_source_newer(source: &Path, output: &Path) -> io::Result<bool> {
    if !output.exists() { return Ok(true); }
    let source_time = fs::metadata(source)?.modified()?;
    let output_time = fs::metadata(output)?.modified()?;
    Ok(source_time > output_time)
}

If a source file has not changed since the last transpilation, the .rs file is not regenerated.


LSP Support (mist-analyzer)

The Mist Language Server Protocol implementation runs a headless rust-analyzer instance and bridges Mist editor requests to Rust positions.

Architecture

Editor (LSP client)
    │
    ▼
┌─────────────────────┐
│   mist-analyzer     │
│   (Rust + Mist)     │
└─────────────────────┘
    │           ▲
    │  JSON-RPC │
    ▼           │
┌─────────────────────┐
│   rust-analyzer     │
│   (headless child)  │
└─────────────────────┘

Flow

  1. Editor sends Mist file edits to mist-analyzer via LSP
  2. mist-analyzer transpiles Mist to Rust
  3. The transpiled Rust is forwarded to the headless rust-analyzer process
  4. For goto-definition, hover, and completion, a unique marker token is injected at the cursor position
  5. The transpiled output with the marker is sent to rust-analyzer
  6. The marker position is located in the response
  7. Results are mapped back to Mist positions using the rev_mapper

Key Features

  • Full document sync — Open, change, save, close
  • Go to definition — Maps through transpiled Rust positions
  • Completions — Supports trigger characters :, ., ', (
  • Hover — Type information and docs
  • Diagnostics — Real-time errors from transpilation failures and Rust compilation
  • Formatting — Document formatting support (currently passthrough)
  • Auto-import — Automatically inserts pub module <name>; when new .mist files are created
  • File watching — Monitors **/*.mist for new files

Diagnostic Remapping

When a transpilation error occurs in the mist-analyzer:

  1. Parse errors are converted to Diagnostic with Mist source positions
  2. Semantic errors (uninitialized class fields) use the stored line/column
  3. Rust compiler errors are remapped via the Mapping system back to Mist positions

Semantic Checks

Field Initialization

The primary semantic check ensures all class fields are initialized in the constructor:

class Player {
    i32 health,
    str& name,

    pub constructor(str& name)
    {
        self.name = name;
        // Error: field 'health' is uninitialized
    }
}

The analysis tracks:

  • Direct field assignment: self.field = value
  • &mut self.field patterns
  • Method calls that might initialize fields (transitively)
  • All conditional branches must initialize the same fields
  • super assignments count for _super field initialization

Position Mapping

Mist maintains a bidirectional mapping between Mist source positions and Rust output positions. This is essential for:

  1. Error remapping — Rust compiler errors point to the correct Mist source location
  2. LSP features — Go-to-definition, hover, and completion work from Mist source

The mapping is stored as HashSet<(RustMap, MistMap)> pairs and serialized to .map.json files alongside the transpiled .rs output.

Mapping Lifecycle

  1. Codegen — Each Spanned<T> AST node records its Mist position and the current Rust position in the codegen output buffer
  2. Persistence — The mapping is written to <output>.map.json
  3. Build-time remapping — The builder reads .map.json to remap cargo diagnostics back to Mist positions
  4. LSP remapping — Both Mist→Rust and Rust→Mist direction queries are supported via find() and find_by_mist()