Zafir Stojanovski
|
832cf46425
|
lint
|
2025-02-12 11:21:46 +01:00 |
|
Zafir Stojanovski
|
52a56cbc4f
|
system prompt for structured output, and parse such outputs
|
2025-02-12 10:44:42 +01:00 |
|
rishabhranawat
|
2d57beb517
|
commit formatting
|
2025-02-10 22:05:45 -08:00 |
|
rishabhranawat
|
6e3d049fed
|
[eval-v1] benchmark with 50 samples
|
2025-02-10 22:05:09 -08:00 |
|
rishabhranawat
|
615c63d2f9
|
[eval-v1] pre commit formatting
|
2025-02-10 21:50:22 -08:00 |
|
rishabhranawat
|
88c875c00f
|
[eval-v1] add timer
|
2025-02-10 21:48:44 -08:00 |
|
rishabhranawat
|
be3d04e7cb
|
[eval-v1] async to speed up inference/evaluation
|
2025-02-10 21:35:46 -08:00 |
|
rishabhranawat
|
03f87dbc07
|
[eval-basic] remove large results files, add gitignore, only leave summary
|
2025-02-09 22:52:10 -08:00 |
|
rishabhranawat
|
2308ed99fb
|
[eval-basic] run precommit formatting
|
2025-02-09 22:40:45 -08:00 |
|
rishabhranawat
|
94f07ed35d
|
[eval-basic] initial scripts for evaluating models on reasoning gym
|
2025-02-09 22:36:27 -08:00 |
|