Inside Predictabl · Case study

Predicting the Next Best Move

We gave Predictabl one sentence. It turned the data into a validated, served prediction model without anyone writing the model code, building the training pipeline or engineering the features by hand.

Mattia●9 min read●September 2026
Predictabl playing Snake using live model predictions

Every turn is a real prediction call. The model runs locally, uses zero output tokens and returns the next move through Predictabl’s API.

Businesses run on decisions.

How much inventory should move to each branch? Which customer should receive an offer? Should this transaction be reviewed? When should a machine be serviced? Some of those decisions are made by people. Others are made automatically by software.

Behind many of them sits a prediction: what is likely to happen next?

The prediction is not the end product. It is an input to an action: something a person can use, a workflow can prioritize, or a system can execute automatically.

Predictions run on tables. LLMs do not.

Most business history lives in rows and columns: transactions, orders, bookings, customers, shifts, machines and events.

LLMs are remarkable with language. They can understand a question, inspect a schema and reason about a sample of data. But they are not good predictive models for millions of tabular records. They are not designed to learn stable patterns across those rows and return a probability you can validate and rely on repeatedly.

Traditional machine learning is very good at exactly that.

The problem is that building a useful ML system has traditionally required much more than training a model. A data scientist still has to understand the question, find the right tables, define the target, prevent leakage, construct features, choose a validation strategy, compare models and make the result usable.

This is where the two technologies fit together.

Predictabl uses an LLM to automate the work around the model, which a data scientist would normally do. It then trains and validates a tabular ML model for the actual prediction.

The LLM helps build the predictive system. It does not make the live prediction.

A small game makes the idea visible

Predicting the Right Next Move — Teaching Predictabl to Play Snake

Snake is a useful toy example because its decisions are easy to see. At every turn, the game has to choose one of three actions: continue straight, turn left or turn right.

In this case, prediction and decision almost coincide. The model predicts the move an expert would make, and the game immediately executes the most likely answer.

We connected the game data to Predictabl and gave it one sentence:

“Predict the right expert move from a single game-state row.”

That sentence was the entire brief.

Nobody wrote the model, the training pipeline or the feature SQL. From that single request, Predictabl worked out what the data meant, built and validated the model, and exposed it as a working prediction API that could play the game.

Predictabl showing the request to predict the right expert move from a raw Snake game-state row
The starting point: connect the data and describe what you want to predict.

The question was simple. The data work was not.

Even for Snake, Predictabl first had to work out what the request meant in data: which table described the game, what one row represented, which column contained the expert’s answer and which fields could accidentally leak it.

The original dataset contained one row per position and 32 signals. Like the feature-assisted representation used by the open-source Laya Snake model, it described whether each possible move would collide, eat, leave enough space or preserve access to the tail. We also included the real path length to the food around the snake’s body.

Starting from that one sentence, Predictabl framed the request as a three-class prediction, selected the relevant columns, excluded the answer-leaking fields, chose a validation strategy and picked LightGBM for the size of the data.

An experienced data scientist could do all of this. But understanding the schema, checking the assumptions and turning the question into a valid training setup would still take hours or days. Predictabl handled the workflow automatically, without hand-written model code or feature SQL.

We also tried removing the hints

Those original columns already contained a lot of useful reasoning. To see how much Predictabl could recover by itself, we ran a second experiment and removed them.

The new dataset contained only a raw 5 × 5 view around the snake’s head, the food and tail positions, its direction and its length. Nothing said “this move is fatal,” “this move eats,” or “this area has enough room.” The cells were stored in fixed board directions, not from the snake’s changing point of view.

We gave Predictabl the same sentence. Nothing more. It rebuilt the board relative to the snake’s direction and reconstructed whether going straight, left or right would be blocked or would reach the food. The collision features never missed a real collision in our checks, and the food features were exact.

Predictabl showing automatically engineered blocked-move and food features from raw Snake board data
From the raw board columns, Predictabl selected 48 signals and engineered 16 additional features without hand-written code or SQL.

Nobody wrote a line of model code or feature SQL. About three minutes later, Predictabl had understood the data, framed the task, built the features, selected the model, trained it and tested it. The raw-board model agreed with the expert on 94.2% of game states it had never seen, with an AUC of 0.989. For comparison, the first model, with the richer precomputed signals, reached an AUC of 0.99996 ± 0.00002 across five folds.

The result is a model, not an LLM answer

Predictabl chose LightGBM for this dataset. It evaluated the model on held-out positions that were not used for fitting, and ran cross-validation before presenting the result.

The important point is not that 94.2% will be sufficient for every application. It is that performance is measured before the model is trusted. Predictabl reports held-out metrics, cross-validation, a confusion matrix and feature importance so a team can decide whether the result is good enough to deploy.

At prediction time, the input is one new raw game-state row. Predictabl applies the saved feature transformation and the trained model returns a probability for each possible move:

straight0.03
left0.94
right0.03

The game takes the move with the highest probability. There is no prompt at this stage, no generated text and no LLM in the loop. It is a prediction call to a small ML model whose behavior was tested before deployment.

The model’s own computation takes roughly one millisecond. In the live demo, a complete call through Predictabl’s local API takes about 60–65 milliseconds, mostly from request overhead, and uses zero output tokens.

Businesses do not store Snake moves

They store transactions, orders, bookings, customer interactions, staffing records, sensor readings and machine events.

But the pattern is the same: connect the historical data and describe the outcome that matters in one sentence. Predictabl works out how that request maps onto the tables, builds the features, chooses and validates a model, and serves a prediction that another person or system can use.

Instead of predicting straight, left or right, the output might be:

Those predictions can inform a person, rank a queue, trigger a workflow or drive an automated decision. They are fast and cheap to call because the live system is a purpose-built ML model, not an LLM guessing from a prompt.

And because the model was evaluated on held-out data before deployment, its quality is explicit. It may or may not be accurate enough for the use case, but that question is answered with evidence rather than confidence-sounding prose.

ML could already do this. The economics are changing.

Machine learning has made decisions from tabular data for years. The model was rarely the expensive part.

The expensive part was everything required to reach it: weeks or months of data discovery, problem framing, feature work, validation and deployment. That is why companies ask a predictive question only when someone can justify a substantial data-science initiative.

Predictabl compresses that surrounding work into an automated path from a plain-language request to a tested, served model.

When asking a predictive question costs almost nothing, the economics of which questions get asked invert.

Prediction no longer needs to be reserved for a few flagship projects. Teams can test many more ideas, keep the ones that prove useful and bring business-specific models into decisions that were previously too small to justify.

Snake is just the visible example. The real opportunity is every useful prediction already hiding in a company’s data.

A small comparison

These systems solve related problems in different ways. The figures below come from different hardware and test setups, so this is a product comparison, not a controlled speed benchmark.

Jev · TypeSafeLaya · open weightsPredictabl
What it isGeneral-purpose decision model with typed outputsOpen-weights decision modelBuilds and serves a task-specific model from your data
How it adaptsNo task-specific training; describe the decisionFine-tuned per task; 41 minutes on a T4 for SnakeTrained automatically on your history; about three minutes here
ModelClosed; architecture undisclosed322M-parameter mmBERT model for SnakeTabular ML; LightGBM here
Surrounding workYou define the inputs, decision contract and evaluationYou prepare the task data, fine-tune and evaluateAutomates framing, feature construction, model choice and validation
DeploymentTypeSafe-hosted APISelf-hosted: local or your cloudPredictabl cloud API or self-hosted; local in this demo
Decision latency236–276 ms p50¹31.8 ms for the fine-tuned Snake model²60–65 ms through the local API; 0.24 ms batched³
Cost modelTrain: none, zero-shot
Inference: paid per decision, grows with volume⁴
Train: you fine-tune each task⁵
Inference: free, on your own compute
Train: under $1 of LLM tokens, once⁶
Inference: $0 in tokens⁷
Learns from your historical outcomesNo, zero-shotOnly if you fine-tune it yourselfYes, that’s the product
Accuracy on your data known before useNoIf you evaluate it yourselfYes, held-out validation built in
Best fitNew decisions with no historyOpen-weight, fine-tuned decision modelsRecurring business predictions with historical data

¹ Independent measurements of Jev over TypeSafe’s hosted API, network round trip included, as reported by third parties; not published by TypeSafe.

² From Laya’s Snake model card, measured on local hardware.

³ Measured on an Apple M3 MacBook Air (16 GB) over localhost, including Predictabl’s per-call API overhead. The model’s own work is well under a millisecond per position; most of the 60–65 ms is per-call overhead (loading the model for each request) that in-memory caching or a production server would largely remove.

⁴ Jev charges $0.042 per million input tokens, and option lists count as input; output is free. Measured on our own requests, one Snake decision is ~120 tokens and one Tetris decision with 34 possible placements is ~720 tokens: about $5–30 per million decisions. Richer business records (a few dozen fields) would plausibly reach 1,000–2,000 tokens, about $40–85 per million. Workloads like fraud checks, pricing or routing can run millions to billions of decisions a year.

⁵ Laya’s Snake model: a full fine-tune of the 322M model, 2 epochs in 41 minutes on a Colab T4 GPU, per its model card.

⁶ Measured for our raw Tetris model: 66,454 LLM tokens (52,434 input, 14,020 output) to explore the data and write the model plan, ≈ $0.60 at list prices, with no LLM calls afterwards. Training took about a minute on a laptop.

⁷ Every prediction is a plain model call, with no LLM involved.

What could your data already predict?

If your company is sitting on years of operational data, Predictabl can help turn it into business-specific predictions you can actually use without starting another months-long data science initiative.

Check out Predictabl ↗