joesharratt1229
|
1a3728ec3a
|
corrected small linting err in cognition.yaml
|
2025-02-11 06:56:04 +00:00 |
|
joesharratt1229
|
bf00437aae
|
converted answer to string
|
2025-02-11 06:48:59 +00:00 |
|
rishabhranawat
|
d7b69190ba
|
commit formatting
|
2025-02-10 22:05:45 -08:00 |
|
rishabhranawat
|
1dc7af587f
|
[eval-v1] benchmark with 50 samples
|
2025-02-10 22:05:09 -08:00 |
|
rishabhranawat
|
fb40c8ca55
|
[eval-v1] add a simple readme with some details
|
2025-02-10 21:57:00 -08:00 |
|
rishabhranawat
|
9e4870125d
|
[eval-v1] pre commit formatting
|
2025-02-10 21:50:22 -08:00 |
|
rishabhranawat
|
df5438498e
|
[eval-v1] add timer
|
2025-02-10 21:48:44 -08:00 |
|
rishabhranawat
|
247464a47d
|
[eval-v1] async to speed up inference/evaluation
|
2025-02-10 21:35:46 -08:00 |
|
joesharratt1229
|
42e02640a3
|
added r1 evaluation logic
|
2025-02-11 03:46:56 +00:00 |
|
rishabhranawat
|
0657222a8f
|
[eval-basic] remove large results files, add gitignore, only leave summary
|
2025-02-09 22:52:10 -08:00 |
|
rishabhranawat
|
c214724a46
|
[eval-basic] run precommit formatting
|
2025-02-09 22:40:45 -08:00 |
|
rishabhranawat
|
75cfd31ec2
|
[eval-basic] initial scripts for evaluating models on reasoning gym
|
2025-02-09 22:36:27 -08:00 |
|