Docs & changelog
This commit is contained in:
parent
86e31b8d29
commit
d4861b6408
3 changed files with 73 additions and 56 deletions
21
CHANGELOG.md
21
CHANGELOG.md
|
|
@ -1,6 +1,27 @@
|
||||||
# Changelog
|
# Changelog
|
||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
|
||||||
|
- **Embeddings**: Migrated to pydantic-ai's embeddings module
|
||||||
|
- Uses pydantic-ai v1.39.0+ embeddings with instrumentation and token counting support
|
||||||
|
- Explicit `embed_query()` and `embed_documents()` API for query/document distinction
|
||||||
|
- New providers available: Cohere (`cohere:`), SentenceTransformers (`sentence-transformers:`)
|
||||||
|
- VoyageAI refactored to extend pydantic-ai's `EmbeddingModel` base class
|
||||||
|
- **Configuration**: Added `base_url` to `ModelConfig` and `EmbeddingModelConfig`
|
||||||
|
- Enables custom endpoints for OpenAI-compatible providers (vLLM, LM Studio, etc.)
|
||||||
|
- Model-level `base_url` takes precedence over provider config
|
||||||
|
|
||||||
|
### Deprecated
|
||||||
|
|
||||||
|
- **vLLM and LM Studio providers**: Use `openai` provider with `base_url` instead
|
||||||
|
- `provider: vllm` → `provider: openai` with `base_url: http://localhost:8000/v1`
|
||||||
|
- `provider: lm_studio` → `provider: openai` with `base_url: http://localhost:1234/v1`
|
||||||
|
|
||||||
|
### Removed
|
||||||
|
|
||||||
|
- Deleted obsolete embedder implementations: `ollama.py`, `openai.py`, `vllm.py`, `lm_studio.py`, `base.py`
|
||||||
|
|
||||||
## [0.22.0] - 2025-12-19
|
## [0.22.0] - 2025-12-19
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
|
|
|
||||||
|
|
@ -135,12 +135,6 @@ providers:
|
||||||
ollama:
|
ollama:
|
||||||
base_url: http://localhost:11434
|
base_url: http://localhost:11434
|
||||||
|
|
||||||
vllm:
|
|
||||||
embeddings_base_url: ""
|
|
||||||
rerank_base_url: ""
|
|
||||||
qa_base_url: ""
|
|
||||||
research_base_url: ""
|
|
||||||
|
|
||||||
docling_serve:
|
docling_serve:
|
||||||
base_url: http://localhost:5001
|
base_url: http://localhost:5001
|
||||||
api_key: ""
|
api_key: ""
|
||||||
|
|
|
||||||
|
|
@ -28,6 +28,7 @@ qa:
|
||||||
- Higher (0.8-1.0+): Creative, varied responses
|
- Higher (0.8-1.0+): Creative, varied responses
|
||||||
- **max_tokens**: Maximum tokens in response
|
- **max_tokens**: Maximum tokens in response
|
||||||
- **enable_thinking**: Control reasoning behavior (see below)
|
- **enable_thinking**: Control reasoning behavior (see below)
|
||||||
|
- **base_url**: Custom endpoint for OpenAI-compatible servers (vLLM, LM Studio, etc.)
|
||||||
|
|
||||||
### Thinking Control
|
### Thinking Control
|
||||||
|
|
||||||
|
|
@ -67,7 +68,7 @@ See the [Pydantic AI thinking documentation](https://ai.pydantic.dev/thinking/)
|
||||||
|
|
||||||
## Embedding Providers
|
## Embedding Providers
|
||||||
|
|
||||||
If you use Ollama, you can use any pulled model that supports embeddings.
|
Embedding models require three settings: `provider`, `name`, and `vector_dim`. Optionally, use `base_url` for OpenAI-compatible servers.
|
||||||
|
|
||||||
### Ollama (Default)
|
### Ollama (Default)
|
||||||
|
|
||||||
|
|
@ -135,41 +136,59 @@ Set your API key via environment variable:
|
||||||
export OPENAI_API_KEY=your-api-key
|
export OPENAI_API_KEY=your-api-key
|
||||||
```
|
```
|
||||||
|
|
||||||
### vLLM
|
### Cohere
|
||||||
|
|
||||||
For high-performance local inference, you can use vLLM to serve embedding models with OpenAI-compatible APIs:
|
Cohere embeddings are available via pydantic-ai:
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
embeddings:
|
embeddings:
|
||||||
model:
|
model:
|
||||||
provider: vllm
|
provider: cohere
|
||||||
|
name: embed-v4.0
|
||||||
|
vector_dim: 1024
|
||||||
|
```
|
||||||
|
|
||||||
|
Set your API key via environment variable:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export CO_API_KEY=your-api-key
|
||||||
|
```
|
||||||
|
|
||||||
|
### SentenceTransformers
|
||||||
|
|
||||||
|
For local embeddings using HuggingFace models:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
embeddings:
|
||||||
|
model:
|
||||||
|
provider: sentence-transformers
|
||||||
|
name: all-MiniLM-L6-v2
|
||||||
|
vector_dim: 384
|
||||||
|
```
|
||||||
|
|
||||||
|
### OpenAI-Compatible Servers (vLLM, LM Studio, etc.)
|
||||||
|
|
||||||
|
For local inference servers with OpenAI-compatible APIs, use the `openai` provider with a custom `base_url`:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
# vLLM example
|
||||||
|
embeddings:
|
||||||
|
model:
|
||||||
|
provider: openai
|
||||||
name: mixedbread-ai/mxbai-embed-large-v1
|
name: mixedbread-ai/mxbai-embed-large-v1
|
||||||
vector_dim: 512
|
vector_dim: 512
|
||||||
|
base_url: http://localhost:8000/v1
|
||||||
|
|
||||||
providers:
|
# LM Studio example
|
||||||
vllm:
|
|
||||||
embeddings_base_url: http://localhost:8000
|
|
||||||
```
|
|
||||||
|
|
||||||
**Note:** You need to run a vLLM server separately with an embedding model loaded.
|
|
||||||
|
|
||||||
### LM Studio
|
|
||||||
|
|
||||||
[LM Studio](https://lmstudio.ai/) provides a local OpenAI-compatible API server for running models:
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
embeddings:
|
embeddings:
|
||||||
model:
|
model:
|
||||||
provider: lm_studio
|
provider: openai
|
||||||
name: text-embedding-qwen3-embedding-4b
|
name: text-embedding-qwen3-embedding-4b
|
||||||
vector_dim: 2560
|
vector_dim: 2560
|
||||||
|
base_url: http://localhost:1234/v1
|
||||||
providers:
|
|
||||||
lm_studio:
|
|
||||||
base_url: http://localhost:1234
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**Note:** LM Studio must be running with an embedding model loaded. The default URL is `http://localhost:1234`.
|
**Note:** The `base_url` must include the `/v1` path for OpenAI-compatible endpoints.
|
||||||
|
|
||||||
## Question Answering Providers
|
## Question Answering Providers
|
||||||
|
|
||||||
|
|
@ -232,45 +251,28 @@ Set your API key via environment variable:
|
||||||
export ANTHROPIC_API_KEY=your-api-key
|
export ANTHROPIC_API_KEY=your-api-key
|
||||||
```
|
```
|
||||||
|
|
||||||
### vLLM
|
### OpenAI-Compatible Servers (vLLM, LM Studio, etc.)
|
||||||
|
|
||||||
For high-performance local inference:
|
For local inference servers with OpenAI-compatible APIs, use the `openai` provider with a custom `base_url`:
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
|
# vLLM example
|
||||||
qa:
|
qa:
|
||||||
model:
|
model:
|
||||||
provider: vllm
|
provider: openai
|
||||||
name: Qwen/Qwen3-4B # Any model with tool support in vLLM
|
name: Qwen/Qwen3-4B
|
||||||
|
base_url: http://localhost:8002/v1
|
||||||
|
|
||||||
providers:
|
# LM Studio example
|
||||||
vllm:
|
|
||||||
qa_base_url: http://localhost:8002
|
|
||||||
```
|
|
||||||
|
|
||||||
**Note:** You need to run a vLLM server separately with a model that supports tool calling loaded. Consult the specific model's documentation for proper vLLM serving configuration.
|
|
||||||
|
|
||||||
### LM Studio
|
|
||||||
|
|
||||||
Use LM Studio for local question answering and research:
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
qa:
|
qa:
|
||||||
model:
|
model:
|
||||||
provider: lm_studio
|
provider: openai
|
||||||
name: openai/gpt-oss-20b
|
name: gpt-oss-20b
|
||||||
|
base_url: http://localhost:1234/v1
|
||||||
enable_thinking: false
|
enable_thinking: false
|
||||||
|
|
||||||
research:
|
|
||||||
model:
|
|
||||||
provider: lm_studio
|
|
||||||
name: openai/gpt-oss-20b
|
|
||||||
|
|
||||||
providers:
|
|
||||||
lm_studio:
|
|
||||||
base_url: http://localhost:1234
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**Note:** LM Studio must be running with a chat model that supports tool calling loaded.
|
**Note:** The server must be running with a model that supports tool calling. The `base_url` must include the `/v1` path.
|
||||||
|
|
||||||
### Other Providers
|
### Other Providers
|
||||||
|
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue