Version 2.5.1#

30.09.26

A new release of Sherpa AI Server has been released: Version 2.5.1.

1. Processing on the Chat Screen#

On the Chat screen, tool calls, their results, and the Assistant's intermediate messages are now grouped in a collapsible block called "Processing." Next to it remains the final answer, and during operation, the block's title shows the current step.

This helps to track the execution of the request and read the final answer without a long chain of service messages.

2. Filters on the Audit Screen#

Server filters by users, time, event type, message, and agent have been added to the Audit screen. They are located in the standard menu of the corresponding columns. Selected values are saved, and export uses the same data area as the table.

The time filter takes into account the account's time zone and does not shift the saved period when the browser switches to daylight saving time. Multiple users and other conditions can now be applied simultaneously.

This helps to check events for the desired period and export only the found records.

3. Changing the Model and Document Index Dimension#

On the Documents screen, in the "Indexing" tab, the saving and reading of the "Indexing Model" and "Embedding Dimension" fields have been fixed. Transitioning between models with vector dimensions of 384 and 1024 is now supported in both directions. When saving a change in model or dimension, old vectors are cleared, and files scheduled for automatic indexing are queued for reprocessing. Samples for search are also reindexed. The manual indexing mode remains manual, and the no-indexing mode does not trigger indexing.

Search takes into account the actual vector dimension and does not mix results from the old and new models. If a file has fragments added only via the API, the system rejects the change of model and separate clearing of vectors: before retrying, such fragments need to be saved outside the system and deleted or the file itself deleted. This protects data that cannot be restored from the original document.

This helps to safely change the indexing model without mixing incompatible search results.

4. Searching Large Document Databases#

For searching in areas exceeding 40 thousand vectors, the use of prepared HNSW indexes with a limited set of candidates has been added. If the number of found active fragments is less than the requested number, the search performs an exact selection in the selected account, files, and dimension. For smaller areas, exact search is used. The administrator can configure the size of the initial candidate set and explicitly select the exact mode via EMBED_SEARCH_STRATEGY=scoped_exact.

This reduces the search workload on large indexes and retrieves the requested number of fragments if they exist in the selected area. The speed and quality of the search depend on the installation data; parameters should be tested on your queries before selection.

5. Sidebar Menu Configuration#

In the client appearance settings file branding.json, the ability to hide the "General" / "Personal" toggle in the sidebar of the Chat screen has been added. The automatic selection of available chats is preserved.

This helps to configure navigation for setups where the user does not need a manual toggle for chat areas.

6. Clearing Chat History with Audit Logging#

In the "Chat Audit Log" section, the "Clear History" action now explicitly warns about the irreversible deletion of messages for the entire account older than the selected date and associated ratings. Deletion and logging of the audit event are performed as a single operation. The event logs the initiator, IP address, cutoff date, and number of deleted messages; for requests with a token, its identifier and name are logged.

This helps the administrator control irreversible clearing and analyze its results through the log.

7. Secure Connection to External Model#

For remote vLLM, an example of configuring a reverse proxy with HTTPS has been added. TLS certificate verification for the language model server can now be enabled in the configuration. Documentation on encrypting media and backups has been supplemented with procedures and limitations for applying AES-256 in the pilot installation.

This helps the administrator plan a secure connection and data storage considering the capabilities of their infrastructure.

8. Preparing to Upgrade a Populated Database#

For an existing database, before starting migrations, build the vector indexes and the duplicate-fragment check index in advance. The scripts backend/migrations/prepare_dimension_hnsw.sql and backend/migrations/prepare_documents_dedup_index.sql are run separately via psql -X -v ON_ERROR_STOP=1 -f, outside the migration transaction. They prepare indexes for dimensions 384 and 1024 without long write blocking. For another dimension, backend/migrations/prepare_additional_dimension_hnsw.sql is required first. The order and parameters are provided in the search documentation.

The vector search service now requires a configured access token when EMBED_AUTH_REQUIRED=true is enabled: a non-empty EMBED_AUTH_TOKEN must be set before starting. The delivery also sets limits on document size and the number of its fragments.

This helps to perform an update on a working database without unexpected lengthy index builds during migration or the first search.

9. Fixes on the Chat Screen#

Autoscroll no longer interrupts reading a long response after manual scrolling up. A smooth scroll down button has been added to return to new messages.

The handling of incoming moderation has been fixed: a fixed response is saved in history as the Assistant's response, and a blocked request does not trigger generation. A false message about exceeding the context of the agent script has been fixed: before truncating the request, the system clarifies its size considering tools and templates. Saved moderation rule sets are loaded again without error.

This helps to read responses at your own pace and receive predictable Assistant behavior under constraints and moderation.

10. Fixes on the Documents Screen#

The document list now loads correctly after refreshing the page, changing the view, and searching. The indicator remains until the request is completed, reloading after an error works, and an outdated response does not replace the new folder filter. After deleting a folder, the filter is updated.

Searching across multiple folders and full-text search BM25 no longer misses existing matches due to premature limiting of the results list. Searching for files without an explicitly specified area is limited to account folders; chat folders are excluded unless the administrator has enabled FILE_SEARCH_INCLUDE_CHAT_FOLDERS=1.

This helps to obtain an up-to-date list of files and expected search results.

11. Document Indexing Fixes#

Indexing via an external model now reports an error if the service response is incomplete, if there are fewer vectors than expected, or if part of the unique fragments is not recorded. In the case of a partial error during clearing, already processed files remain in the reindexing queue. Matching text from different files or with different origins is saved separately.

An empty status field when uploading a file via the API is now equivalent to a missing field: the file receives the standard pending status. During manual indexing, the API accepts an empty body {} and determines the account from the authorization.

This helps to avoid counting incomplete indexing as successful and maintains the connection of search results with the correct source.

12. Prompt Library and User Session#

In the prompt library, opening folders and switching between the "Personal" / "General" tabs has been accelerated: the list loads page by page without a separate request for each folder. Long API requests no longer hold the user session and do not interfere with other requests. The session lifetime now corresponds to SESSION_LIFETIME_MINUTES, so the PHP garbage collector does not prematurely terminate it after 24 minutes of inactivity.

This makes working with the library and parallel actions in the interface more responsive.

13. API and Server Services Fixes#

Incorrect scalar values for the filters object_folder_guids and field_names now undergo regular data validation instead of a PHP error. An error in the field name file_guids instead of files_guids in the search query returns a 422 code. The OpenAPI validation reuses the uploaded schema within the request.

The timeout for Whisper is now set via WHISPER_UPSTREAM_TIMEOUT_SECONDS; by default, it is 310 seconds. In the client nginx, the timeout for long requests has been increased to 600 seconds. The PostgreSQL password has been excluded from the indexing service logs.

This simplifies the diagnosis of integration errors and reduces the likelihood of interruptions during long processing.

14. Reliability of Updates and Operation of Assistants#

Migrations that could stop due to a repeating version number, a long extension of an old file, or connecting to PostgreSQL via a local socket have been fixed. After removing a model field, the metadata cache is cleared before creating a Doctrine proxy. The updated version of pg_textsearch eliminates indexing failures after a BM25 index error.

On the Assistants screen, editing, pinning, and deleting are again determined by permissions for actions with assistants: there is no longer a separate setting to prohibit editing. The duplication of the sign in confirmation dialogs has also been fixed.

This reduces the risk of failures during updates and returns control of assistants to the assigned role permissions.