Small Language Model Benchmark
Pick a model and run the LAMBADA word-prediction task through the OpenRouter API. Results are saved to the results folder and shown below.
Hyperparameters
These control decoding, not model weights. Optimal uses greedy decoding with the most worked examples (most reliable accuracy), Normal is the balanced default, and Best Performance trims the token and example budget for the fastest, cheapest runs.
lambada — benchmark output
Metrics
| Model | Developer | Params | Accuracy (%) | Correct | Total | Avg time (s) | Errors |
|---|