Overview
The WDS server is built from eight .Web service projects. It can run as one in-process service or as a distributed service set.
Service inventory
| Service | Responsibility |
|---|---|
| Dapi | Public REST, Swagger, and MCP gateway; job persistence and orchestration. |
| Crawler | HTTP downloads, request controls, and Crawler address registration. |
| Datakeeper | Job settings, download-task state, downloaded content, and cache storage. |
| Scraper | Structured extraction, content conversion, and scraped-data storage. |
| Idealer | Stable identifier allocation and tenant cleanup. |
| Retriever | Full-text and vector indexing and search. |
| Jober | Traversal configuration, execution, scheduling, recovery, and results. |
| Solidstack | Single-container implementation of the public gateway and backend capabilities. |
Docs and Playground are auxiliary deployment containers, not WDS server services. Compose and Helm deployments may include them for offline documentation and repeatable examples; see Auxiliary containers.
Single-service mode
Solidstack runs the Dapi, Crawler, Datakeeper, Idealer, Jober, Scraper, and Retriever implementations in one process. It is suited to evaluation and small workloads, but its components cannot be scaled or restarted independently. The current Solidstack feature provider does not enable Scheduling. See Plans for feature availability.
Multi-service mode
Multi-service deployments run Dapi, Crawler, Datakeeper, Scraper, Idealer, Retriever, and Jober as separate processes. This topology supports independent scaling, health monitoring, resource allocation, and failure isolation. It is available with the Business plan.
Dapi is the public entry point. It stores job configurations in MongoDB and calls the other services over gRPC-Web. Crawler registers with Datakeeper Resource Manager; Datakeeper, Scraper, and Retriever use Idealer; and Jober coordinates Dapi, Datakeeper, and Scraper for Traversal runs.
Runtime dependencies
- MongoDB stores service state. Every distributed service except Crawler uses it; Solidstack uses one MongoDB configuration for its in-process components.
- Datakeeper can use a separate MongoDB or S3-compatible cache. Without one, it uses its primary MongoDB database.
- Retriever can use a separate MongoDB Atlas Search database. Vector modes also require an HTTP embedding service.
- Dapi, Crawler, Datakeeper, Idealer, Jober, Scraper, and Retriever require service-specific license configuration in multi-service mode.