Commit graph

18 commits

Author SHA1 Message Date
Yiorgis Gozadinos
87f28ed96b
Update docs 2025-09-30 11:30:35 +03:00
Yiorgis Gozadinos
7cc4e13b57
Make benchmarks use pydantic evals 2025-09-30 11:30:34 +03:00
Yiorgis Gozadinos
a215a64686
Refactor to give its research subagent full responsibility 2025-09-19 12:10:00 +03:00
Yiorgis Gozadinos
d90dfe258c
Use logfire when running evaluations 2025-09-12 09:45:06 +03:00
Yiorgis Gozadinos
2207b7d1ed
vacuum() in store, client, cli. Removes old versions and reduces disk usage 2025-09-10 11:24:51 +03:00
Yiorgis Gozadinos
3c235e8b6c
Suppress logging from deps 2025-09-03 11:42:08 +03:00
Yiorgis Gozadinos
b73949c784
update benchmarks 2025-09-02 15:10:57 +03:00
Yiorgis Gozadinos
27b16bdf8d
Update basic benchmarks 2025-09-02 10:51:20 +03:00
Yiorgis Gozadinos
b05e5f5408
Set/get haiku version to db 2025-09-01 17:41:41 +03:00
Yiorgis Gozadinos
628bdc8151
Cleanup 2025-09-01 14:56:15 +03:00
Yiorgis Gozadinos
b83596ff07
Basic moving to lancedb. 2025-09-01 14:56:14 +03:00
Yiorgis Gozadinos
42e62f92a1
Better formatting of benchmark script 2025-07-09 10:27:32 +03:00
Yiorgis Gozadinos
9d72a158a6
Document benchmarks 2025-07-08 19:01:54 +03:00
Yiorgis Gozadinos
1705e7a21c
Honour QA_MODEL for anthropic and OpenAI 2025-07-04 12:14:13 +03:00
Yiorgis Gozadinos
85b106c461
OpenAI Question/Answer agent 2025-06-28 09:11:39 +03:00
Yiorgis Gozadinos
ca45f8e5d4
Fix db path 2025-06-25 19:01:57 +03:00
Yiorgis Gozadinos
2855dd7c13
Use LLM-as-a-judge to test QA 2025-06-25 18:18:28 +03:00
Yiorgis Gozadinos
26900d9687
Run simple benchmarks 2025-06-25 18:18:28 +03:00