Andreas Köpf
d344284c80
feat: Add comprehensive visualization script for evaluation results ( #376 )
2025-03-16 12:26:27 +01:00
Andreas Köpf
677a2af03e
Add eval configs, small fixes to eval script & rush-hour score_answer
2025-03-16 09:18:05 +01:00
Andreas Koepf
fa950d0189
add gemma-3-27b & qwq-32b configs
2025-03-15 20:47:51 +01:00
Andreas Köpf
424ee6751a
Eval N completions per prompt ( #374 )
...
* feat: Add support for generating multiple completions per prompt
* feat: Track best and mean scores for multiple completions per prompt
* feat: Add checkpoint and resume functionality to evaluation script
2025-03-15 16:39:36 +01:00
joesharratt1229
1f6de829bd
Algebra/curr ( #320 )
...
* add polynomial equation curriculum
* added simple integration
* addded metadata to config
2025-03-11 00:17:07 +01:00
Andreas Koepf
1d813c9acd
update eval yaml config files
2025-03-10 00:48:32 +01:00
Andreas Köpf
0b108efac1
Generate eval config tool ( #240 )
...
* feat: Add generate_config.py script to create eval configurations
2025-02-27 21:40:53 +01:00
Andreas Köpf
1ea9a657a7
Eval script consolidation ( #238 )
...
The script now supports:
- YAML and JSON configurations
- Dataset-specific parameters
- Overriding configuration via command line
- Detailed logging and error handling
2025-02-27 17:39:14 +01:00
Andreas Koepf
726ba114dc
add llama-3.3-70b-instruct eval yaml files
2025-02-26 20:54:07 +01:00
Andreas Köpf
c0e5941fe5
Merge pull request #217 from open-thought/feat/o3-mini-eun
...
added o3 mini yaml rconfiguration
2025-02-26 09:38:11 +01:00
vncntt
98af865309
fix sonnet eval_dir ( #216 )
...
* fix eval_dir
* add logging
2025-02-26 09:37:09 +01:00
joesharratt1229
8eaece6f05
added o3 mini yaml
2025-02-26 08:09:12 +00:00
Andreas Koepf
9b7eec2d64
add llama-3.3-70b-instruct algebra, algorithmic eval configs
2025-02-25 23:43:29 +01:00
Andreas Koepf
a60cdb0775
use results folder name for eval results
2025-02-25 19:41:21 +01:00
joesharratt1229
e0e8bab09c
Merge remote-tracking branch 'origin/consolidate_eval_script' into fix/eval
2025-02-25 18:10:07 +00:00
joesharratt1229
68e8ea89d8
changed structure
2025-02-25 16:32:42 +00:00
joesharratt1229
93b95d748b
updated read me
2025-02-25 15:46:43 +00:00
Andreas Koepf
11fb7e0edf
move r1 configs into r1 yaml/r1 subfolder
2025-02-25 16:24:30 +01:00
Andreas Koepf
7f0047667f
consolidate eval scripts to have single eval.py
2025-02-25 16:13:22 +01:00