Commit graph

57 commits

Author SHA1 Message Date
Andreas Köpf
c8e77d21a7
fix: Improve error logging and preserve full model response in eval process (#337) 2025-03-12 00:01:49 +01:00
Andreas Koepf
770255b608 fix pre-commit 2025-03-11 08:18:55 +01:00
joesharratt1229
105374183f
Algebra/curr (#320)
* add polynomial equation curriculum

* added simple integration

* addded metadata to config
2025-03-11 00:17:07 +01:00
Andreas Koepf (aider)
d0b49cfffd feat: Add --category option to evaluate datasets from a specific category 2025-03-11 00:00:38 +01:00
Andreas Koepf
4109b5b72c update eval yaml config files 2025-03-10 00:48:32 +01:00
vncntt
3672b231f1
should exit if API key isn't defined (#259)
* should exit if open-router and no api key
2025-03-04 09:45:36 +01:00
joesharratt1229
6770ee3eef
updated for config by dataset (#257)
* updated for config by dataset

* updated read me
2025-03-03 21:58:32 +01:00
Andreas Köpf
a66a7e7965
Revert "log error message on bad api response (#243)" (#249)
This reverts commit 8e2089b6c0.
2025-03-01 23:56:42 +01:00
Andreas Köpf
4ad9d22fa3
Add base_url and api_key command line args for eval.py script (#244)
* feat: Add base URL command line parameter to eval.py script
* feat: Add API key parameter and CLI option to AsyncModelEvaluator
2025-02-28 18:32:58 +01:00
Rich Jones
8e2089b6c0
log error message on bad api response (#243) 2025-02-28 15:32:27 +01:00
Andreas Köpf
b4207162ff
Eval sampling settings for generation (temperature, top-p, max_tokens) (#242)
* feat: Add sampling parameters to eval configuration and API call
* feat: Add support for system_prompt_id and optional system_prompt configuration
2025-02-28 11:48:37 +01:00
Andreas Koepf (aider)
24a4b7a4c8 feat: Add system prompt to dataset results and summary output 2025-02-28 00:26:06 +01:00
Andreas Köpf
5b8d1b5175
Generate eval config tool (#240)
* feat: Add generate_config.py script to create eval  configurations
2025-02-27 21:40:53 +01:00
Andreas Köpf
850c1cf6f4
Eval script consolidation (#238)
The script now supports:
   - YAML and JSON configurations
   - Dataset-specific parameters
   - Overriding configuration via command line
   - Detailed logging and error handling
2025-02-27 17:39:14 +01:00
Andreas Koepf
477e1f85cc verify that OPENROUTER_API_KEY env var is set 2025-02-26 22:15:30 +01:00
Andreas Koepf
acb2d7eb53 add llama-3.3-70b-instruct eval yaml files 2025-02-26 20:54:07 +01:00
Andreas Köpf
5b89a3a2d0
Merge pull request #217 from open-thought/feat/o3-mini-eun
added o3 mini yaml rconfiguration
2025-02-26 09:38:11 +01:00
vncntt
29179f783e
fix sonnet eval_dir (#216)
* fix eval_dir

* add logging
2025-02-26 09:37:09 +01:00
joesharratt1229
7d7e44d1af added o3 mini yaml 2025-02-26 08:09:12 +00:00
Andreas Koepf
6d5168d1e5 add llama-3.3-70b-instruct algebra, algorithmic eval configs 2025-02-25 23:43:29 +01:00
Andreas Koepf
791f16ec0f use results folder name for eval results 2025-02-25 19:41:21 +01:00
joesharratt1229
ffe60ef112 finalised readme 2025-02-25 18:14:39 +00:00
joesharratt1229
56cc111ab3 Merge remote-tracking branch 'origin/consolidate_eval_script' into fix/eval 2025-02-25 18:10:07 +00:00
joesharratt1229
9ac6ea4eb2 changed structure 2025-02-25 16:32:42 +00:00
joesharratt1229
52c3c430b9 updated config and read me 2025-02-25 16:25:16 +00:00
joesharratt1229
7b39f4a3c7 updated read me 2025-02-25 15:51:31 +00:00
joesharratt1229
046c46c0bb updated read me 2025-02-25 15:46:43 +00:00
Andreas Koepf
878f9bbc76 move r1 configs into r1 yaml/r1 subfolder 2025-02-25 16:24:30 +01:00
Andreas Koepf
e7ae82a831 consolidate eval scripts to have single eval.py 2025-02-25 16:13:22 +01:00
Andreas Koepf
8291956554 add aiohttp & tenacity deps to requirements-eval.txt 2025-02-25 15:50:11 +01:00
Andreas Koepf (aider)
e48c1f82cd docs: Update installation instructions in eval README 2025-02-25 15:37:09 +01:00
Andreas Koepf (aider)
a1b0a0414e docs: Add dependency installation step to eval README setup instructions 2025-02-25 15:19:38 +01:00
Andreas Koepf
574edb5c5b remove eval results from main repo 2025-02-25 11:02:02 +01:00
Andreas Koepf (aider)
205174c532 docs: Add info about reasoning-gym-eval repository for evaluation results 2025-02-25 10:53:21 +01:00
joesharratt1229
cffbff935c pinned provider to nebius 2025-02-24 05:01:22 +00:00
Andreas Köpf
2947038557
Merge pull request #182 from zafstojano/env/binary-alternation
feat(env): Binary Alternation
2025-02-21 17:27:16 +01:00
Andreas Koepf
3e7ff3b084 use native types List->list, Dict->dict, Set->set, Tuple->tuple 2025-02-21 15:15:38 +01:00
Zafir Stojanovski
77789257d3 include pre-parsed responses in json 2025-02-21 13:50:48 +01:00
Zafir Stojanovski
d557b1b4f9 contribution updates 2025-02-20 09:54:26 +01:00
Andreas Köpf
e14e61824f
Merge pull request #123 from joesharratt1229/feat/r1-evals
Added r1 async implementation and algorithmic config
2025-02-13 11:37:07 +01:00
joesharratt1229
b2e3ccf3d6 updated async impl and added r1 2025-02-13 03:51:01 +00:00
Zafir Stojanovski
58a641e59f lint 2025-02-12 11:21:46 +01:00
Zafir Stojanovski
7c43b8ee47 reset eval.sh 2025-02-12 10:58:49 +01:00
Zafir Stojanovski
3d84816f95 system prompt for structured output, and parse such outputs 2025-02-12 10:44:42 +01:00
Andreas Köpf
858e5bba51
Merge pull request #108 from rishabhranawat/eval-v2
Eval V1: improve speed using async
2025-02-11 16:07:47 +01:00
joesharratt1229
1a3728ec3a corrected small linting err in cognition.yaml 2025-02-11 06:56:04 +00:00
joesharratt1229
bf00437aae converted answer to string 2025-02-11 06:48:59 +00:00
rishabhranawat
d7b69190ba commit formatting 2025-02-10 22:05:45 -08:00
rishabhranawat
1dc7af587f [eval-v1] benchmark with 50 samples 2025-02-10 22:05:09 -08:00
rishabhranawat
fb40c8ca55 [eval-v1] add a simple readme with some details 2025-02-10 21:57:00 -08:00