reasoning-gym

mirror of https://github.com/open-thought/reasoning-gym.git synced 2026-04-19 12:58:07 +00:00

Author	SHA1	Message	Date
joesharratt1229	2eea347296	updated read me	2025-02-25 15:51:31 +00:00
joesharratt1229	93b95d748b	updated read me	2025-02-25 15:46:43 +00:00
Andreas Koepf	4eb4933647	add aiohttp & tenacity deps to requirements-eval.txt	2025-02-25 15:50:11 +01:00
Andreas Koepf (aider)	795685f30e	docs: Update installation instructions in eval README	2025-02-25 15:37:09 +01:00
Andreas Koepf (aider)	bb0d1f0a82	docs: Add dependency installation step to eval README setup instructions	2025-02-25 15:19:38 +01:00
Andreas Koepf	d40da704db	remove eval results from main repo	2025-02-25 11:02:02 +01:00
Andreas Koepf (aider)	a073a2792b	docs: Add info about reasoning-gym-eval repository for evaluation results	2025-02-25 10:53:21 +01:00
joesharratt1229	1b0f774974	pinned provider to nebius	2025-02-24 05:01:22 +00:00
Andreas Köpf	de362fb76f	Merge pull request #182 from zafstojano/env/binary-alternation feat(env): Binary Alternation	2025-02-21 17:27:16 +01:00
Andreas Koepf	ff5b210106	use native types List->list, Dict->dict, Set->set, Tuple->tuple	2025-02-21 15:15:38 +01:00
Zafir Stojanovski	0391a99446	include pre-parsed responses in json	2025-02-21 13:50:48 +01:00
Zafir Stojanovski	4b9a874a61	contribution updates	2025-02-20 09:54:26 +01:00
Andreas Köpf	00d0e7a57d	Merge pull request #123 from joesharratt1229/feat/r1-evals Added r1 async implementation and algorithmic config	2025-02-13 11:37:07 +01:00
joesharratt1229	abf74c7e7a	updated async impl and added r1	2025-02-13 03:51:01 +00:00
Zafir Stojanovski	832cf46425	lint	2025-02-12 11:21:46 +01:00
Zafir Stojanovski	80f4ae8457	reset `eval.sh`	2025-02-12 10:58:49 +01:00
Zafir Stojanovski	52a56cbc4f	system prompt for structured output, and parse such outputs	2025-02-12 10:44:42 +01:00
Andreas Köpf	05ec556ede	Merge pull request #108 from rishabhranawat/eval-v2 Eval V1: improve speed using async	2025-02-11 16:07:47 +01:00
joesharratt1229	ecddc3aa9f	corrected small linting err in cognition.yaml	2025-02-11 06:56:04 +00:00
joesharratt1229	9df5f45083	converted answer to string	2025-02-11 06:48:59 +00:00
rishabhranawat	2d57beb517	commit formatting	2025-02-10 22:05:45 -08:00
rishabhranawat	6e3d049fed	[eval-v1] benchmark with 50 samples	2025-02-10 22:05:09 -08:00
rishabhranawat	06cabcfdee	[eval-v1] add a simple readme with some details	2025-02-10 21:57:00 -08:00
rishabhranawat	615c63d2f9	[eval-v1] pre commit formatting	2025-02-10 21:50:22 -08:00
rishabhranawat	88c875c00f	[eval-v1] add timer	2025-02-10 21:48:44 -08:00
rishabhranawat	be3d04e7cb	[eval-v1] async to speed up inference/evaluation	2025-02-10 21:35:46 -08:00
joesharratt1229	a3ea4449d1	added r1 evaluation logic	2025-02-11 03:46:56 +00:00
rishabhranawat	03f87dbc07	[eval-basic] remove large results files, add gitignore, only leave summary	2025-02-09 22:52:10 -08:00
rishabhranawat	2308ed99fb	[eval-basic] run precommit formatting	2025-02-09 22:40:45 -08:00
rishabhranawat	94f07ed35d	[eval-basic] initial scripts for evaluating models on reasoning gym	2025-02-09 22:36:27 -08:00

30 commits