joesharratt1229
|
2eea347296
|
updated read me
|
2025-02-25 15:51:31 +00:00 |
|
joesharratt1229
|
93b95d748b
|
updated read me
|
2025-02-25 15:46:43 +00:00 |
|
Andreas Koepf
|
4eb4933647
|
add aiohttp & tenacity deps to requirements-eval.txt
|
2025-02-25 15:50:11 +01:00 |
|
Andreas Koepf (aider)
|
795685f30e
|
docs: Update installation instructions in eval README
|
2025-02-25 15:37:09 +01:00 |
|
Andreas Koepf (aider)
|
bb0d1f0a82
|
docs: Add dependency installation step to eval README setup instructions
|
2025-02-25 15:19:38 +01:00 |
|
Andreas Koepf
|
d40da704db
|
remove eval results from main repo
|
2025-02-25 11:02:02 +01:00 |
|
Andreas Koepf (aider)
|
a073a2792b
|
docs: Add info about reasoning-gym-eval repository for evaluation results
|
2025-02-25 10:53:21 +01:00 |
|
joesharratt1229
|
1b0f774974
|
pinned provider to nebius
|
2025-02-24 05:01:22 +00:00 |
|
Andreas Köpf
|
de362fb76f
|
Merge pull request #182 from zafstojano/env/binary-alternation
feat(env): Binary Alternation
|
2025-02-21 17:27:16 +01:00 |
|
Andreas Koepf
|
ff5b210106
|
use native types List->list, Dict->dict, Set->set, Tuple->tuple
|
2025-02-21 15:15:38 +01:00 |
|
Zafir Stojanovski
|
0391a99446
|
include pre-parsed responses in json
|
2025-02-21 13:50:48 +01:00 |
|
Zafir Stojanovski
|
4b9a874a61
|
contribution updates
|
2025-02-20 09:54:26 +01:00 |
|
Andreas Köpf
|
00d0e7a57d
|
Merge pull request #123 from joesharratt1229/feat/r1-evals
Added r1 async implementation and algorithmic config
|
2025-02-13 11:37:07 +01:00 |
|
joesharratt1229
|
abf74c7e7a
|
updated async impl and added r1
|
2025-02-13 03:51:01 +00:00 |
|
Zafir Stojanovski
|
832cf46425
|
lint
|
2025-02-12 11:21:46 +01:00 |
|
Zafir Stojanovski
|
80f4ae8457
|
reset eval.sh
|
2025-02-12 10:58:49 +01:00 |
|
Zafir Stojanovski
|
52a56cbc4f
|
system prompt for structured output, and parse such outputs
|
2025-02-12 10:44:42 +01:00 |
|
Andreas Köpf
|
05ec556ede
|
Merge pull request #108 from rishabhranawat/eval-v2
Eval V1: improve speed using async
|
2025-02-11 16:07:47 +01:00 |
|
joesharratt1229
|
ecddc3aa9f
|
corrected small linting err in cognition.yaml
|
2025-02-11 06:56:04 +00:00 |
|
joesharratt1229
|
9df5f45083
|
converted answer to string
|
2025-02-11 06:48:59 +00:00 |
|
rishabhranawat
|
2d57beb517
|
commit formatting
|
2025-02-10 22:05:45 -08:00 |
|
rishabhranawat
|
6e3d049fed
|
[eval-v1] benchmark with 50 samples
|
2025-02-10 22:05:09 -08:00 |
|
rishabhranawat
|
06cabcfdee
|
[eval-v1] add a simple readme with some details
|
2025-02-10 21:57:00 -08:00 |
|
rishabhranawat
|
615c63d2f9
|
[eval-v1] pre commit formatting
|
2025-02-10 21:50:22 -08:00 |
|
rishabhranawat
|
88c875c00f
|
[eval-v1] add timer
|
2025-02-10 21:48:44 -08:00 |
|
rishabhranawat
|
be3d04e7cb
|
[eval-v1] async to speed up inference/evaluation
|
2025-02-10 21:35:46 -08:00 |
|
joesharratt1229
|
a3ea4449d1
|
added r1 evaluation logic
|
2025-02-11 03:46:56 +00:00 |
|
rishabhranawat
|
03f87dbc07
|
[eval-basic] remove large results files, add gitignore, only leave summary
|
2025-02-09 22:52:10 -08:00 |
|
rishabhranawat
|
2308ed99fb
|
[eval-basic] run precommit formatting
|
2025-02-09 22:40:45 -08:00 |
|
rishabhranawat
|
94f07ed35d
|
[eval-basic] initial scripts for evaluating models on reasoning gym
|
2025-02-09 22:36:27 -08:00 |
|