Every `Store` built its own connection with its own caches and discarded them on close, so the index a vector query loads was refetched by the next connection. On object storage that first fetch dominates: measured on a ~500k-chunk 2560-dim corpus over a ~200ms link, the first query cost ~41s and the second ~3s, and a new connection reusing the session cost ~7s instead of ~47s. `connect_lancedb` now passes a process-wide session, keyed on the configured cache sizes so a caller asking for different sizes gets its own. Also sets `read_consistency_interval`, defaulting to 30s. It was None, meaning a connection never re-checked for other processes' writes. Per-call connections hid that; a shared session makes connections long-lived enough for a reader to go stale against the ingester. All three settings reject negatives at the config boundary. A negative cache size raises OverflowError and a negative interval panics inside Lance, so neither is catchable further in. Zero stays valid for both: no cache, and check on every read. The routing tests now assert the kwargs they care about rather than the full call signature, since every connection carries the two new kwargs. |
||
|---|---|---|
| .. | ||
| capabilities | ||
| configuration | ||
| img | ||
| stylesheets | ||
| apps.md | ||
| benchmarks.md | ||
| changelog.md | ||
| chat.md | ||
| cli.md | ||
| custom-pipelines.md | ||
| development.md | ||
| index.md | ||
| ingester.md | ||
| installation.md | ||
| mcp.md | ||
| overview.md | ||
| python.md | ||
| remote-processing.md | ||
| tools.md | ||
| tuning.md | ||
| tutorial.md | ||