Yiorgis Gozadinos
fd327996b8
Configurable judge and reflect models for evaluations
2026-03-18 17:14:11 +02:00
Yiorgis Gozadinos
6ac5f277d8
Rename iterations to num_candidates
2026-03-12 13:00:51 +02:00
Yiorgis Gozadinos
24275d2ef9
Fixes
2026-03-12 12:19:21 +02:00
Yiorgis Gozadinos
c785a54f5f
Respect config.prompts.qa in evaluations benchmark and optimization
2026-03-12 12:05:51 +02:00
Yiorgis Gozadinos
6812435ba0
cleanup
2026-03-12 12:05:50 +02:00
Yiorgis Gozadinos
8bac8b4104
Properly choose training/eval set
2026-03-12 12:05:50 +02:00
Yiorgis Gozadinos
45f7ac8d1b
run_optimization should be sync
2026-03-12 12:05:25 +02:00
Yiorgis Gozadinos
60b1d3c013
Add GEPA prompt optimization for QA evaluations
2026-03-12 12:05:25 +02:00
Yiorgis Gozadinos
e2749ad2a6
Cap QA agent search iterations to reduce response time
2026-03-11 12:13:43 +02:00
Yiorgis Gozadinos
40ff54eac0
Set params for evaluations judge
2026-03-05 13:59:20 +02:00
Yiorgis Gozadinos
4855051936
Add research() to client and add haiku.skills dependency, remove --deep flag and simplify app
2026-02-20 16:36:47 +02:00
Yiorgis Gozadinos
42f2d0fd0e
docs
2026-01-31 10:13:22 +02:00
Yiorgis Gozadinos
2a99cb09bf
Evaluate using ask --deep
2026-01-29 14:22:18 +02:00
Yiorgis Gozadinos
39f4937916
Handle missing datasets from huggingface
2026-01-29 13:31:31 +02:00
Yiorgis Gozadinos
369b3a4bf7
No zip
2026-01-26 16:28:19 +02:00
Yiorgis Gozadinos
9b166969e2
Introduce huggingface dataset for sharing evaluation dbs
2026-01-26 15:53:37 +02:00
Yiorgis Gozadinos
c7a9ad8583
Customize orb QA prompt to not use LaTeX as gpt-oss Ollama implementation fails to parse it properly
2026-01-24 13:48:09 +02:00
Yiorgis Gozadinos
de022639b6
Correctly lookup docs for the orb dataset
2026-01-22 14:17:09 +02:00
Yiorgis Gozadinos
e6b239bd6c
Add evaluation dataset for multi-modal q&a
2026-01-22 14:17:09 +02:00
Yiorgis Gozadinos
ea7a4a9a8f
apply dotenv usecwd fix to remaining files
...
Apply the same find_dotenv(usecwd=True) fix from cli.py to:
- evaluations/evaluations/benchmark.py
- app/backend/main.py
Co-Authored-By: Tianyi Cui <contact@tianyicui.com>
2026-01-22 12:41:35 +02:00
Yiorgis Gozadinos
7cc561d1db
Refactor to make agents a top-level module. Bring in the conversational agent from the app
2026-01-12 12:37:08 +02:00
Yiorgis Gozadinos
40e0ea090d
Add option to override vacuum interval
2025-12-12 16:04:06 +02:00
Yiorgis Gozadinos
056b9ad090
Use periodic vacuum to prevent disk exhaustion with large datasets
2025-12-12 15:53:23 +02:00
Yiorgis Gozadinos
802f058205
Move search related config under SearchConfig
2025-12-09 12:28:28 +02:00
Yiorgis Gozadinos
06b4f6a224
Add format parameter for text-to-DoclingDocument conversion
2025-12-08 15:56:21 +02:00
Yiorgis Gozadinos
9a5f8e4368
Add type-aware context expansion for search results
2025-12-08 15:56:21 +02:00
Yiorgis Gozadinos
d808c6c425
Simplify citations in qa & graph agents
2025-12-08 15:55:27 +02:00
Yiorgis Gozadinos
ace58a7034
Update evaluations and a2a
2025-12-08 15:54:56 +02:00
Yiorgis Gozadinos
3525fae625
Use EmbeddingModelConfig similar to ModelConfig for embeddings
2025-12-02 11:55:31 +02:00
Yiorgis Gozadinos
8a8005ead0
Add qa/judge models meta to experiment meta
2025-11-25 13:09:29 +02:00
Yiorgis Gozadinos
f843fb8140
Use non-thinking judge in evals
2025-11-25 13:00:25 +02:00
Yiorgis Gozadinos
e83978f88b
Revert embeddings to use flat config
2025-11-25 12:51:58 +02:00
Yiorgis Gozadinos
a9701616c8
Rename model to name under model
2025-11-25 12:28:43 +02:00
Yiorgis Gozadinos
1bb7b6d5bd
Add support for per-model configuration settings including thinking, temperature and max_tokens
2025-11-25 12:14:24 +02:00
Yiorgis Gozadinos
4e1752a740
Default eval db location, evaluation script
2025-11-25 11:07:20 +02:00
Yiorgis Gozadinos
b5892699a0
Use Mean Reciprocal Rank for single document evaluation as metric. Use Mean Average Precision for variable document evaluation as metric
2025-11-24 14:49:30 +02:00
Yiorgis Gozadinos
8d0846b88a
Record useful meta in the experiment
2025-11-24 10:46:25 +02:00
Yiorgis Gozadinos
63b374c7e9
Name evaluation runs. Use only LLMJudge evaluator.
2025-11-21 13:26:39 +02:00
Yiorgis Gozadinos
adb7d9093d
Run entire evaluation dataset as one so that it appears properly in logfire
2025-11-14 11:39:58 +02:00
Yiorgis Gozadinos
f55b8c8f19
Load .env in evals
2025-11-14 11:00:58 +02:00
Yiorgis Gozadinos
b8bc6d9b68
Let ruff know about our package structure
2025-11-05 17:47:27 +02:00
Yiorgis Gozadinos
009e529869
No need to have evaluations in the docker image
2025-11-05 12:35:37 +02:00
Yiorgis Gozadinos
2f9c907031
Restructure into uv workspace to support minimal and full installations
2025-11-04 17:59:12 +02:00