Commit graph

29 commits

Author SHA1 Message Date
Andreas Köpf
677a2af03e Add eval configs, small fixes to eval script & rush-hour score_answer 2025-03-16 09:18:05 +01:00
Andreas Koepf (aider)
a829707408 feat: Add fallback to first non-None model answer when best_answer is unset 2025-03-15 16:52:50 +01:00
Andreas Köpf
424ee6751a Eval N completions per prompt (#374)
* feat: Add support for generating multiple completions per prompt
* feat: Track best and mean scores for multiple completions per prompt
* feat: Add checkpoint and resume functionality to evaluation script
2025-03-15 16:39:36 +01:00
Andreas Köpf
db688ce884 fix: Improve error logging and preserve full model response in eval process (#337) 2025-03-12 00:01:49 +01:00
Andreas Koepf
59e309b732 fix pre-commit 2025-03-11 08:18:55 +01:00
Andreas Koepf (aider)
d29a665081 feat: Add --category option to evaluate datasets from a specific category 2025-03-11 00:00:38 +01:00
vncntt
478646622e should exit if API key isn't defined (#259)
* should exit if open-router and no api key
2025-03-04 09:45:36 +01:00
Andreas Köpf
16a4ea1193 Revert "log error message on bad api response (#243)" (#249)
This reverts commit 27e66ba6dd.
2025-03-01 23:56:42 +01:00
Andreas Köpf
dbd2ac723e Add base_url and api_key command line args for eval.py script (#244)
* feat: Add base URL command line parameter to eval.py script
* feat: Add API key parameter and CLI option to AsyncModelEvaluator
2025-02-28 18:32:58 +01:00
Rich Jones
27e66ba6dd log error message on bad api response (#243) 2025-02-28 15:32:27 +01:00
Andreas Köpf
59922486c6 Eval sampling settings for generation (temperature, top-p, max_tokens) (#242)
* feat: Add sampling parameters to eval configuration and API call
* feat: Add support for system_prompt_id and optional system_prompt configuration
2025-02-28 11:48:37 +01:00
Andreas Koepf (aider)
82e79d672e feat: Add system prompt to dataset results and summary output 2025-02-28 00:26:06 +01:00
Andreas Köpf
0b108efac1 Generate eval config tool (#240)
* feat: Add generate_config.py script to create eval  configurations
2025-02-27 21:40:53 +01:00
Andreas Köpf
1ea9a657a7 Eval script consolidation (#238)
The script now supports:
   - YAML and JSON configurations
   - Dataset-specific parameters
   - Overriding configuration via command line
   - Detailed logging and error handling
2025-02-27 17:39:14 +01:00
Andreas Koepf
4cd5bd42c3 verify that OPENROUTER_API_KEY env var is set 2025-02-26 22:15:30 +01:00
vncntt
98af865309 fix sonnet eval_dir (#216)
* fix eval_dir

* add logging
2025-02-26 09:37:09 +01:00
Andreas Koepf
9b7eec2d64 add llama-3.3-70b-instruct algebra, algorithmic eval configs 2025-02-25 23:43:29 +01:00
joesharratt1229
e0e8bab09c Merge remote-tracking branch 'origin/consolidate_eval_script' into fix/eval 2025-02-25 18:10:07 +00:00
joesharratt1229
93b95d748b updated read me 2025-02-25 15:46:43 +00:00
Andreas Koepf
11fb7e0edf move r1 configs into r1 yaml/r1 subfolder 2025-02-25 16:24:30 +01:00
Andreas Koepf
7f0047667f consolidate eval scripts to have single eval.py 2025-02-25 16:13:22 +01:00
Andreas Köpf
de362fb76f Merge pull request #182 from zafstojano/env/binary-alternation
feat(env): Binary Alternation
2025-02-21 17:27:16 +01:00
Andreas Koepf
ff5b210106 use native types List->list, Dict->dict, Set->set, Tuple->tuple 2025-02-21 15:15:38 +01:00
Zafir Stojanovski
0391a99446 include pre-parsed responses in json 2025-02-21 13:50:48 +01:00
Zafir Stojanovski
52a56cbc4f system prompt for structured output, and parse such outputs 2025-02-12 10:44:42 +01:00
rishabhranawat
615c63d2f9 [eval-v1] pre commit formatting 2025-02-10 21:50:22 -08:00
rishabhranawat
88c875c00f [eval-v1] add timer 2025-02-10 21:48:44 -08:00
rishabhranawat
be3d04e7cb [eval-v1] async to speed up inference/evaluation 2025-02-10 21:35:46 -08:00
rishabhranawat
03f87dbc07 [eval-basic] remove large results files, add gitignore, only leave summary 2025-02-09 22:52:10 -08:00
Renamed from eval/eval_basic.py (Browse further)