Chess Trainer

GitHub ยท Website

I’m going to be really honest: I’m a bad chess player, I’ve always found chess interesting, but never really learned to play it well.

But as they say, “Those who can, do; those who can’t, teach”, so I wanted to see if I could give a chess model some advice to follow, and maybe even make it play better by giving the right advice.

That’s why “Chess Trainer” exists: start with a naive Maia model, give your player some advice, then play against it or try someone else’s. First access gives you a private account link, where you can access to your players without logging, while an optional email lets you recover it.

How does this work?

Vanilla Maia

Every player uses Maia-3 5M, a transformer that predicts human moves given a rating: I chose 900, which means that the model will predict the moves a 900 ELO player would do.

The transformer processes board history and ratings, producing a 256-dimensional representation for each of the 64 squares, then Maia’s move head turns these into move scores, and finally a wrapper restricts the choice to legal moves and picks the highest-scoring one.

Teaching Maia

An LLM1 interprets the advice, then the request is formalized as both a preference and a scorer that can score how much each move adhere to the advice.

Then using Python I train a small neural network on examples selected from a bank of 50k positions from human games, with precomputed activations for Maia@900 and Maia@2500: the 2500 rating pass adds is used to add context, and both use the same exact weights, the only difference is in the input I provide.

The network returns a 256-dimensional delta vector for each square, added to Maia’s transformer output before the move head, on top of an existing frozen prudent control: training will favour moves that score well under the advice among reasonable alternatives, while limiting changes where the advice is irrelevant.

Maia’s weights are not touched in any way and will stay frozen, what changes is the vector that gets added between the transformer and the move head.

Combining advice

To combine multiple rules, I tested something similar to one step of gradient boosting: train an adapter for each rule, freeze these single adapters, then train an additional corrector for that specific combination.

All adapters, including the corrector, read the original Maia activations at 900 and 2500, then the runtime sums their deltas before the move head.

This does not work everytime though, and in some cases the adapter is a single one, since it gets a better adherence to the requested rules.

Be aware that the adapter does not apply strict rules or constraints, just steers the model towards preferring moves that actually comply with the rules.

How to play

Go to the website, create your set of rules, create your player and just unleash it in the competition! Or play against it, to check how strong they are!


  1. Qwen3.8 27B 4bit running on my homelab, since it’s already there ↩︎