Change the document indexing model#
This guide explains how to change the model that creates vectors for document search. When you change the model or vector dimension, Sherpa AI Server clears the old vectors and indexes the documents again. Content search may be unavailable during the switch, and results may be incomplete until file processing finishes. Schedule the change for a suitable time.
Before you start#
- You need the
documents:updateandai_models:readpermissions to change these settings. Creating a model or model instance also requiresai_models:create. Configuring a vLLM instance through the API requiresai_models:update. - Find the new model's actual vector dimension. Enter the number it returns: an integer from 32 to 2000. When upgrading from 2.5.0 to 2.5.1, prepare the 384 and 1024 indexes. If your current version is earlier than 2.5.0, agree the upgrade path with support first. For any other dimension, an administrator must prepare the search index first.
- A model of type Indexing or Multimodal must have exactly one active instance. Multiple instances of the same model may return incompatible vectors.
- Do not change the address or path of the current indexing model's active instance to point it to another vector source: the system rejects that change. Create a separate model and instance, then select the new model in the document indexing settings.
- If documents contain chunks added only through the API, without a source file, read “If the change is rejected” below. Ordinary reindexing cannot restore these chunks.
Example: BGE-M3 with vLLM#
On a GPU server, start a separate vLLM instance for the indexing model. This example uses vLLM 0.27.1 and BAAI/bge-m3 (vLLM must be installed and the model files available):
vllm serve BAAI/bge-m3 --runner pooling --pooler-config.task embed \
--hf-overrides '{"architectures":["BgeM3EmbeddingModel"]}' \
--host 0.0.0.0 --port 8001
On the vLLM server, check that the model returns 1024 numbers per vector (jq is required):
curl -fsS http://127.0.0.1:8001/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{"model":"BAAI/bge-m3","input":["test"]}' \
| jq '.data[0].embedding | length'
In Models, create a new Indexing model and one active local instance. Set the model path to BAAI/bge-m3, the port to 8001, and the vLLM address to one that is reachable from the Sherpa AI Server container. If vLLM runs on another server, use its network name or IP address. 127.0.0.1 inside the Sherpa container refers to that container itself. Leave the API subpath empty.
Before selecting this model for indexing, configure its instance through the Sherpa AI Server API with ai_models:update permission: call GET /api/v1/models/{model_guid}/instances/{instance_guid}, copy the entire data.connection_params object, add "embed_format": "openai", and send the complete updated object in a PATCH to the same URL as {"connection_params": {…}}. Call GET again and confirm that data.connection_params.embed_format is openai. A PATCH replaces the whole connection parameters object, so retain every existing connection and authentication field.
This setting routes Sherpa requests to /v1/embeddings. Without it, a local instance uses /v2/embed with an input_type value that BGE-M3 rejects in vLLM 0.27.1. After configuring the instance, do not save it from the Models form: the form does not expose embed_format and removes it on save. The form's connection test only checks server availability; Sherpa checks the actual vectors when you save the indexing settings. Then follow the steps below to select the new model and dimension 1024.
Change the model#
- Open Models and check that the required model and its single active instance are available. If the model is missing and you have
ai_models:create, select New in the model list, create a model of type Indexing or Multimodal, then select New in the instance list and add an instance. Otherwise, contact an administrator. - Open Documents → Settings → Indexing. Select the Indexing model and enter its actual Embedding dimension. Changing the number alone does not change the model.
- Select Save once. The server checks the model, the dimension it returns, and whether the search index is ready. After a successful check, it saves the new model and dimension, clears the old vectors, and queues documents configured for automatic indexing. The original files are preserved. You do not need to select Delete all vectors and reindex separately. Selecting a different model also starts the change when its dimension is unchanged.
- Wait for document processing: Pending → Processing → Ready. For a file with Awaiting manual indexing status, open it and select Re-index, then wait for Ready. Files with Not indexed status are not indexed. The server reindexes assistant answer samples automatically.
- Search for content in a known document and confirm that the required files have reached Ready. If a file shows Error, read and resolve the cause. Then select Re-index for that file. To process all files again, use Delete all vectors and reindex separately.
A transition status of ready means the settings change has finished; it does not mean every file has reached Ready. If an API response shows queued_files, it counts files placed in the queue, not files already processed. Check document statuses and search after indexing finishes.
Examples#
- 384 → 1024. Select a new model that returns 1024-dimensional vectors, enter 1024, and select Save. Wait for file reindexing and check search.
- 1024 → 384. Select a model that returns 384-dimensional vectors, enter 384, and follow the same steps.
- 384 → 384, different model. Select a different model, keep 384, and select Save. Changing the model starts reindexing even though the dimension stays the same.
If the change is rejected or interrupted#
If the preliminary check fails because the model is unavailable, the dimension is wrong, or the index has not been prepared, resolve the cause and save again. If the change is rejected because of chunks added only through the API, the previous model, dimension, and vectors are preserved. Save the text and metadata for these chunks outside the system, then ask an administrator to remove them or delete their file if it is no longer needed. After the model change succeeds, add the saved chunks through the API again if required.
If the transition has started but remains pending or is interrupted, reopen the settings and save the same model and dimension. The server completes cleanup and queues files; the indexing itself runs in the background. Do not select another model and dimension pair while the transition is incomplete.
If you also changed PDF recognition settings, chunk size, or chunk overlap, reopen the settings after an error and check those fields: they may have been saved separately. This does not mean the model and dimension changed successfully.
For administrators: another dimension#
For a dimension other than 384 or 1024, prepare the index for the account after the version 2.5.1 migrations and before saving the new model. The full SQL, checks, and commands for the standard client installation are in the “Additional dimensions after the upgrade” section of the index preparation guide. Copy the SQL from there to a local file and specify the numeric account ID and actual model dimension. After the command succeeds, the user can select Save.