cwiesen
a68cb23b6e
fix: resolve incorrect dataset and column namings
2026-08-14 16:29:51 -05:00
cwiesen
93c21272d1
feat: add search_filter to evaluations
2026-08-14 16:29:51 -05:00
Yiorgis Gozadinos
e522d8cfa8
Remove the wix evaluation dataset
2026-08-14 14:46:51 +03:00
Yiorgis Gozadinos
f4d8e744a2
Record this release's ORB multimodal numbers
...
The nemotron rows for retrieval and for both capabilities are re-measured on
this release with no reranker, judged by Qwen3.6-35B with thinking on. Rows for
other embedders and older versions keep their own attribution.
2026-08-13 13:00:02 +03:00
Yiorgis Gozadinos
ba963864f3
Pin the eval judge sampling and standardise on Qwen3-Reranker
2026-08-06 13:17:58 +03:00
Yiorgis Gozadinos
9deb1f2bd4
replace haiku.skills with native Pydantic AI capabilities
2026-07-24 15:26:17 +03:00
Yiorgis Gozadinos
a72c94aa77
Add ORB multimodal-reranker retrieval result to benchmarks
2026-07-24 13:25:07 +03:00
Yiorgis Gozadinos
3dd234ecfc
Add reranked hotpotqa benchmark results.
2026-07-19 12:08:16 +03:00
Yiorgis Gozadinos
11303aa714
Document hotpotqa benchmark results and finalize the reference config
2026-07-17 16:28:16 +03:00
Yiorgis Gozadinos
3cb229d2e0
Add t2_finqa pre-built evaluation database reference config and docs
2026-06-29 10:52:23 +03:00
Yiorgis Gozadinos
4d03b1e669
Add nemotron-vl multimodal eval database and reference configs
2026-06-29 10:14:53 +03:00
Yiorgis Gozadinos
9ef25e2e53
Order benchmarks docs ORB, T², Wix
2026-06-08 11:51:21 +03:00
Yiorgis Gozadinos
a3a73f1331
Add --filter-ids to run QA on a case-id subset
2026-06-08 09:36:01 +03:00
Yiorgis Gozadinos
4ff17a84bb
Update benchmarks
2026-06-02 15:20:17 +03:00
Yiorgis Gozadinos
cbd9b46da6
Results for analysis skill
2026-06-01 10:40:52 +03:00
Yiorgis Gozadinos
e653c49fce
Remove unecessary Mean Reciprocal Rank metric
2026-06-01 10:40:52 +03:00
Yiorgis Gozadinos
00f2a40a60
Refresh benchmarks doc and remove unused eval datasets
2026-06-01 10:40:51 +03:00
Yiorgis Gozadinos
6ac7d5a4be
Add nemotron embedder to ORB benchmarks and consolidate source buckets
2026-06-01 10:40:51 +03:00
Yiorgis Gozadinos
e881ce285d
update orb benchmark
2026-05-21 14:23:41 +03:00
Yiorgis Gozadinos
0ee00334c7
drop prose emdashes and semicolons; minor fixes
2026-05-20 14:27:16 +03:00
Yiorgis Gozadinos
f1f2e24639
docs: benchmarks — skill-only framing, move inactive datasets to bottom
2026-05-20 14:27:16 +03:00
Yiorgis Gozadinos
910d178b00
docs: rebuild Skills section, drop pitch prose
2026-05-20 14:27:16 +03:00
Yiorgis Gozadinos
6f95e2bc27
Delete the standalone QA and analysis agents
2026-05-19 11:39:20 +03:00
Yiorgis Gozadinos
f588e1150b
add vllm Gemma-4-26B skill QA row to wix table
2026-05-18 09:43:13 +03:00
Yiorgis Gozadinos
abc5a210a6
update benchmarks
2026-05-15 10:04:17 +03:00
Yiorgis Gozadinos
d76c645862
Update benchmarks for orb for gemma4
2026-05-14 14:03:19 +03:00
Yiorgis Gozadinos
2189002974
Benchmark update
2026-05-12 17:20:48 +03:00
Yiorgis Gozadinos
cc59dc7203
Use upload_large_folder for hf upload
2026-05-06 13:12:08 +03:00
Yiorgis Gozadinos
ce7201271f
split open_rag_bench dataset into orb_text and orb_multimodal variants
2026-05-06 12:49:03 +03:00
Yiorgis Gozadinos
a45d82f206
Update benchmark
2026-05-05 13:58:48 +03:00
Yiorgis Gozadinos
5c3a864a0b
rename open_rag_bench eval DBs to text/multimodal variants
2026-05-05 12:55:06 +03:00
Yiorgis Gozadinos
17e36fc069
Add first openrag results
2026-05-05 09:50:22 +03:00
Yiorgis Gozadinos
d509b0662d
update benchmarks
2026-04-30 16:54:17 +03:00
Yiorgis Gozadinos
3391335e3a
Update qa wix benchmarks
2026-04-29 14:21:40 +03:00
Yiorgis Gozadinos
318266e681
Update benchmark
2026-04-29 13:27:32 +03:00
Yiorgis Gozadinos
8962987ff9
pin judge model to ollama:qwen3.6
2026-04-29 12:24:06 +03:00
Yiorgis Gozadinos
5e31c15907
docs
2026-04-28 14:44:27 +03:00
Yiorgis Gozadinos
fd327996b8
Configurable judge and reflect models for evaluations
2026-03-18 17:14:11 +02:00
Yiorgis Gozadinos
37f28ea8de
Clean up stale references and dead code
2026-02-20 16:37:15 +02:00
Yiorgis Gozadinos
a889883cfc
Update wix benchmarks
2026-02-10 12:49:24 +02:00
Yiorgis Gozadinos
5f6488e116
Update benchmarks
2026-01-30 13:05:47 +02:00
Yiorgis Gozadinos
2a99cb09bf
Evaluate using ask --deep
2026-01-29 14:22:18 +02:00
Yiorgis Gozadinos
369b3a4bf7
No zip
2026-01-26 16:28:19 +02:00
Yiorgis Gozadinos
9b166969e2
Introduce huggingface dataset for sharing evaluation dbs
2026-01-26 15:53:37 +02:00
Yiorgis Gozadinos
50a1a6171c
Update orb qa accuracy
2026-01-26 10:54:36 +02:00
Yiorgis Gozadinos
756285e91c
Add retrieval benchmarks for ORB
2026-01-22 15:01:37 +02:00
Yiorgis Gozadinos
9bf3a83b5d
Fix typo
2026-01-05 16:39:36 +02:00
Yiorgis Gozadinos
d77bf89aad
Update retrieval/qa benchmark for hotpotqa
2025-12-17 08:27:40 +02:00
Yiorgis Gozadinos
f726229f60
Reformat benchmarks and add hotpotqa placeholder
2025-12-12 12:33:31 +02:00
Yiorgis Gozadinos
be0e3c3472
Update benchmarks
2025-12-10 12:04:07 +02:00