reasoning-gym

mirror of https://github.com/open-thought/reasoning-gym.git synced 2026-04-22 16:49:06 +00:00

Author	SHA1	Message	Date
Andreas Köpf	677a2af03e	Add eval configs, small fixes to eval script & rush-hour score_answer	2025-03-16 09:18:05 +01:00
Andreas Koepf	fa950d0189	add gemma-3-27b & qwq-32b configs	2025-03-15 20:47:51 +01:00
Andreas Koepf (aider)	a829707408	feat: Add fallback to first non-None model answer when best_answer is unset	2025-03-15 16:52:50 +01:00
Andreas Köpf	424ee6751a	Eval N completions per prompt (#374 ) * feat: Add support for generating multiple completions per prompt * feat: Track best and mean scores for multiple completions per prompt * feat: Add checkpoint and resume functionality to evaluation script	2025-03-15 16:39:36 +01:00
Andreas Köpf	db688ce884	fix: Improve error logging and preserve full model response in eval process (#337 )	2025-03-12 00:01:49 +01:00
Andreas Koepf	59e309b732	fix pre-commit	2025-03-11 08:18:55 +01:00
joesharratt1229	1f6de829bd	Algebra/curr (#320 ) * add polynomial equation curriculum * added simple integration * addded metadata to config	2025-03-11 00:17:07 +01:00
Andreas Koepf (aider)	d29a665081	feat: Add --category option to evaluate datasets from a specific category	2025-03-11 00:00:38 +01:00
Andreas Koepf	1d813c9acd	update eval yaml config files	2025-03-10 00:48:32 +01:00
vncntt	478646622e	should exit if API key isn't defined (#259 ) * should exit if open-router and no api key	2025-03-04 09:45:36 +01:00
joesharratt1229	340d6a7ab9	updated for config by dataset (#257 ) * updated for config by dataset * updated read me	2025-03-03 21:58:32 +01:00
Andreas Köpf	16a4ea1193	Revert "log error message on bad api response (#243 )" (#249 ) This reverts commit `27e66ba6dd`.	2025-03-01 23:56:42 +01:00
Andreas Köpf	dbd2ac723e	Add base_url and api_key command line args for eval.py script (#244 ) * feat: Add base URL command line parameter to eval.py script * feat: Add API key parameter and CLI option to AsyncModelEvaluator	2025-02-28 18:32:58 +01:00
Rich Jones	27e66ba6dd	log error message on bad api response (#243 )	2025-02-28 15:32:27 +01:00
Andreas Köpf	59922486c6	Eval sampling settings for generation (temperature, top-p, max_tokens) (#242 ) * feat: Add sampling parameters to eval configuration and API call * feat: Add support for system_prompt_id and optional system_prompt configuration	2025-02-28 11:48:37 +01:00
Andreas Koepf (aider)	82e79d672e	feat: Add system prompt to dataset results and summary output	2025-02-28 00:26:06 +01:00
Andreas Köpf	0b108efac1	Generate eval config tool (#240 ) * feat: Add generate_config.py script to create eval configurations	2025-02-27 21:40:53 +01:00
Andreas Köpf	1ea9a657a7	Eval script consolidation (#238 ) The script now supports: - YAML and JSON configurations - Dataset-specific parameters - Overriding configuration via command line - Detailed logging and error handling	2025-02-27 17:39:14 +01:00
Andreas Koepf	4cd5bd42c3	verify that OPENROUTER_API_KEY env var is set	2025-02-26 22:15:30 +01:00
Andreas Koepf	726ba114dc	add llama-3.3-70b-instruct eval yaml files	2025-02-26 20:54:07 +01:00
Andreas Köpf	c0e5941fe5	Merge pull request #217 from open-thought/feat/o3-mini-eun added o3 mini yaml rconfiguration	2025-02-26 09:38:11 +01:00
vncntt	98af865309	fix sonnet eval_dir (#216 ) * fix eval_dir * add logging	2025-02-26 09:37:09 +01:00
joesharratt1229	8eaece6f05	added o3 mini yaml	2025-02-26 08:09:12 +00:00
Andreas Koepf	9b7eec2d64	add llama-3.3-70b-instruct algebra, algorithmic eval configs	2025-02-25 23:43:29 +01:00
Andreas Koepf	a60cdb0775	use results folder name for eval results	2025-02-25 19:41:21 +01:00
joesharratt1229	3a2de98b1c	finalised readme	2025-02-25 18:14:39 +00:00
joesharratt1229	e0e8bab09c	Merge remote-tracking branch 'origin/consolidate_eval_script' into fix/eval	2025-02-25 18:10:07 +00:00
joesharratt1229	68e8ea89d8	changed structure	2025-02-25 16:32:42 +00:00
joesharratt1229	ce8877167d	updated config and read me	2025-02-25 16:25:16 +00:00
joesharratt1229	2eea347296	updated read me	2025-02-25 15:51:31 +00:00
joesharratt1229	93b95d748b	updated read me	2025-02-25 15:46:43 +00:00
Andreas Koepf	11fb7e0edf	move r1 configs into r1 yaml/r1 subfolder	2025-02-25 16:24:30 +01:00
Andreas Koepf	7f0047667f	consolidate eval scripts to have single eval.py	2025-02-25 16:13:22 +01:00
Andreas Koepf	4eb4933647	add aiohttp & tenacity deps to requirements-eval.txt	2025-02-25 15:50:11 +01:00
Andreas Koepf (aider)	795685f30e	docs: Update installation instructions in eval README	2025-02-25 15:37:09 +01:00
Andreas Koepf (aider)	bb0d1f0a82	docs: Add dependency installation step to eval README setup instructions	2025-02-25 15:19:38 +01:00
Andreas Koepf	d40da704db	remove eval results from main repo	2025-02-25 11:02:02 +01:00
Andreas Koepf (aider)	a073a2792b	docs: Add info about reasoning-gym-eval repository for evaluation results	2025-02-25 10:53:21 +01:00
joesharratt1229	1b0f774974	pinned provider to nebius	2025-02-24 05:01:22 +00:00
Andreas Köpf	de362fb76f	Merge pull request #182 from zafstojano/env/binary-alternation feat(env): Binary Alternation	2025-02-21 17:27:16 +01:00
Andreas Koepf	ff5b210106	use native types List->list, Dict->dict, Set->set, Tuple->tuple	2025-02-21 15:15:38 +01:00
Zafir Stojanovski	0391a99446	include pre-parsed responses in json	2025-02-21 13:50:48 +01:00
Zafir Stojanovski	4b9a874a61	contribution updates	2025-02-20 09:54:26 +01:00
Andreas Köpf	00d0e7a57d	Merge pull request #123 from joesharratt1229/feat/r1-evals Added r1 async implementation and algorithmic config	2025-02-13 11:37:07 +01:00
joesharratt1229	abf74c7e7a	updated async impl and added r1	2025-02-13 03:51:01 +00:00
Zafir Stojanovski	832cf46425	lint	2025-02-12 11:21:46 +01:00
Zafir Stojanovski	80f4ae8457	reset `eval.sh`	2025-02-12 10:58:49 +01:00
Zafir Stojanovski	52a56cbc4f	system prompt for structured output, and parse such outputs	2025-02-12 10:44:42 +01:00
Andreas Köpf	05ec556ede	Merge pull request #108 from rishabhranawat/eval-v2 Eval V1: improve speed using async	2025-02-11 16:07:47 +01:00
joesharratt1229	ecddc3aa9f	corrected small linting err in cognition.yaml	2025-02-11 06:56:04 +00:00

1 2

61 commits