Benchmarks
Benchmarks & Leaderboards
Benchmark results from the Tiki AI team. When a weakness we fix for a client is common across the industry, the eval behind it hardens into a public benchmark.
Latest Publication
BenchmarkSeptember 2026
Assistant Reliability Bench
An evaluation of persistent, tool-driven agent workflows: 28 long-horizon tasks across 10 workflow categories, scored with deterministic checks, rubric assessment, and repeated-run analysis.
AgentsTool UseLong-HorizonReliability
LeaderboardTotal / 100
- 01Claude Fable 5.186.45
- 02GPT-6 Astra85.38
- 03Kimi K378.23
- 04DeepSeek V4.1 Flash76.69
- 05GLM-5.375.35
View full results
BenchmarkJune 2026
Image-to-Code: A Dataset & Benchmark for Static Web Page Replication
A dataset and benchmark for turning web page screenshots into code, scored on 8 metrics validated against human review. Fine-tuning on 1,000 samples lifts the base model 16.5 points.
Image-to-CodeMultimodalDatasetFine-tuning
LeaderboardMean score
- 01GPT-5.480.02
- 02Qwen-3.6-plus72.64
- 03Doubao-seed-2.072.10
- 04Kimi-k2.670.58
View full results
What's Next
More on the way
New benchmarks will be added here as our team publishes them, each with a leaderboard and the full method behind its scores.