Search...
Ctrl K
Models
Providers
Apps
Rankings
Playground
Models
Providers
Apps
Rankings
Playground
Search...
Ctrl K
Sign In
Sign In
GSM8K Chat - Benchmark Leaderboard & Model Performance | AI Stats
GSM8K Chat
Overview
Overview
Type: percentage
Math
Recorded Results
1
Average Score
81.88
Score Range
81.88 - 81.88
Leading Model
81.88 - Llama 3.1 Nemotron 70B Instruct
Scores Over Time
Individual benchmark scores plotted by date.
Models Using This Benchmark
Organisation
Model
Reported
Top Score
Info
Self Reported
Source
Nvidia
Llama 3.1 Nemotron 70B Instruct
01 Oct 2024
81.88
-
Yes
Source