Victus Infrastructure provides the shared runtime environment used by the Victus ecosystem.
It defines the common services, networking, persistence, observability, secret management, and deployment automation required by other Victus systems.
The infrastructure repository does not contain application logic, scientific processing logic, retrieval behavior, or conversational reasoning. Its responsibility is to provide stable runtime capabilities that other systems can consume through explicit endpoints and contracts.
Victus Infrastructure is responsible for:
Victus Infrastructure does not own:
The production runtime is organized into four infrastructure stacks.
Each stack has a distinct responsibility and can be deployed independently while sharing selected runtime networks.
The current stacks are:
core
observability
llm
wiki
Docker Compose is the source of truth for runtime topology.
Ansible prepares the host and deploys the Compose-defined runtime.
The Core stack contains the infrastructure services intended to be shared by Victus systems.
Its main components are:
SeaweedFS
PostgreSQL
Redis
etcd
CoreDNS
private NGINX
public NGINX
Conceptually:
Victus Systems
↓
Private DNS / Edge
↓
Shared Runtime Services
↓
PostgreSQL / Redis / S3
The Core stack provides persistence, event transport, object storage, service discovery, and network routing.
Victus Infrastructure provides several forms of persistent state.
A shared PostgreSQL service is available for systems that require durable registry or coordination state.
Infrastructure owns the runtime service.
Consumer repositories own their application schemas and migrations.
This means:
victus-infra
→ owns PostgreSQL availability
consumer repository
→ owns database schema and domain semantics
The infrastructure layer should not become the owner of application tables.
SeaweedFS provides an S3-compatible storage interface.
It is intended for:
Object storage is shared infrastructure, but the meaning and lifecycle of domain-specific artifacts remain owned by the producing system.
Redis is configured with persistence and Redis Streams support.
Its primary role is operational event transport.
Redis should not be considered the final source of truth for durable application state.
Conceptually:
PostgreSQL
→ durable state
Redis Streams
→ operational event delivery
Consumers are responsible for handling duplicate events, acknowledgements, retries, and recovery against their authoritative durable state.
Persistent service data is stored outside the lifecycle of individual containers.
This allows services to survive container restarts and redeployments.
Current persistence includes data for:
This protects against container replacement but does not currently protect against complete VPS or disk loss.
Victus networking separates private infrastructure traffic from explicitly public services.
Tailscale provides the primary private access layer for infrastructure services.
Services intended only for Victus systems or operators should remain accessible through the private network rather than being directly exposed to the public Internet.
CoreDNS provides private DNS for shared Victus infrastructure.
The private zone is:
victus.io
Examples of infrastructure service names include:
s3.victus.io
pipeline-postgres.victus.io
redis.victus.io
litellm.victus.io
langfuse.victus.io
Consumer systems should prefer stable infrastructure DNS names instead of depending on Docker container names or host paths.
Private HTTP services such as LiteLLM, Langfuse, and S3 access can be routed through NGINX bound to the private network.
This provides a stable service boundary without exposing internal container ports directly.
Public HTTP access is explicitly controlled.
The current primary public infrastructure service is Wiki.js.
Internet
↓
Public NGINX
↓
Wiki.js
Public NGINX exposes HTTP/HTTPS and uses Let's Encrypt certificates.
Other infrastructure services should not become public by default.
Victus Infrastructure provides a dedicated LLM runtime stack.
The main components are:
LiteLLM
Langfuse
LLM PostgreSQL
LiteLLM provides a shared OpenAI-compatible model gateway.
Its role is to centralize model-provider access and related gateway behavior.
Consumer systems can use the gateway without needing to know deployment details of the LLM stack.
Langfuse provides LLM-focused tracing, cost inspection, and audit visibility.
It is separate from infrastructure-level metrics and logging.
The LLM stack stores persistent application state in its own PostgreSQL runtime.
Provider credentials and runtime secrets are not committed to Git.
Infrastructure observability currently has two main layers.
Infrastructure
→ Prometheus
→ Loki
→ Alloy
LLM activity
→ LiteLLM
→ Langfuse
Prometheus stores infrastructure metrics.
Loki stores aggregated logs.
Alloy is configured on the host as the telemetry collection layer.
Langfuse provides observability specifically for LLM requests and traces.
The current observability system provides the foundations for metrics and logs but is not yet operationally complete.
Missing capabilities currently include:
Observability should therefore be considered implemented but partial.
Wiki.js is hosted as its own infrastructure stack.
It uses a dedicated PostgreSQL database and is exposed through the public edge.
Wiki.js is the central documentation system for Victus.
Repository documentation should focus on implementation and operation, while architecture and ecosystem knowledge lives in Wiki.js.
The infrastructure repository owns the runtime hosting Wiki.js, not the architectural content stored inside it.
Production secrets are managed outside Git.
The current flow uses Infisical as the central secret manager.
GitHub Actions obtains secrets through OIDC and stages the required runtime configuration during deployment.
Conceptually:
Runtime secret files are written on the host with restricted permissions.
The repository contains configuration templates and secret expectations, but not production credentials.
Production deployment is automated through GitHub Actions and Ansible.
The deployment flow is:
The deployment process performs:
The current stack deployment order is approximately:
observability
↓
core
↓
llm
↓
wiki
Docker Compose remains the runtime source of truth.
Ansible coordinates deployment but should not duplicate service topology.
The same Compose-defined topology is used as the basis for local and production execution.
Conceptually:
Local
→ Compose + local configuration
Production
→ Compose + production configuration
→ deployed through Ansible
This keeps local validation relevant to production without requiring the VPS to become the primary development environment.
Some first-time environment preparation remains manual.
Examples include:
Victus Infrastructure exposes shared services that other systems may consume.
Conceptually:
The infrastructure repository defines these shared capabilities, but it does not automatically prove that every Victus system is actively consuming them.
Consumer integrations must be verified in the corresponding application repositories.
Current infrastructure documentation should therefore distinguish:
available infrastructure
≠
confirmed active consumer
Shared infrastructure contracts should remain infrastructure-focused.
Examples of valid infrastructure guarantees include:
Domain semantics should remain outside Infrastructure.
For example:
Infrastructure owns:
S3 service and bucket availability
Scientific Processing owns:
meaning and lifecycle of scientific artifacts
Similarly:
Infrastructure owns:
Redis Streams runtime
Consumer system owns:
event names, payload meaning, retry behavior
This prevents Infrastructure from becoming the accidental owner of application-domain contracts.
The current infrastructure provides reasonable durability against container restarts and ordinary redeployments.
It does not yet provide strong disaster recovery.
Current limitations include:
The current reliability model can therefore be summarized as:
Container restart / redeploy
→ reasonably protected
Host disk corruption / VPS loss
→ weak protection
Backup and restore capability is the most important missing infrastructure reliability feature.
A future production baseline should include:
Automated Backup
↓
Offsite Copy
↓
Retention Policy
↓
Restore Procedure
↓
Periodic Restore Test
Public HTTPS is handled through NGINX and Let's Encrypt / Certbot.
Certificate issuance is integrated into deployment.
However, continuous certificate renewal independently from deployment has not yet been demonstrated as part of the current runtime.
Certificate lifecycle management should eventually be verified independently from application deployments.
The current infrastructure provides a substantial shared runtime foundation.
Implemented and repository-validated capabilities include:
Partially implemented or operationally unverified capabilities include:
Not currently implemented include:
The repository should therefore be considered a well-defined shared runtime foundation with incomplete operational hardening and disaster recovery.
Service topology belongs in Docker Compose.
Deployment automation should deploy that topology rather than recreate it independently.
Infrastructure provides runtime capabilities.
Consumer systems own domain behavior, schemas, workers, and application logic.
Shared infrastructure services should remain private unless public exposure is explicitly required.
Consumers should depend on stable infrastructure endpoints and DNS names rather than internal container or host implementation details.
Production credentials are obtained from external secret management and staged only at runtime.
Replacing or redeploying containers should not destroy durable state.
Validation and deployment should be scripted and repeatable rather than dependent on undocumented manual steps.
A service being deployed by Infrastructure does not mean a consumer repository actively uses it.
Persistence alone is not sufficient. Backup, offsite retention, and tested restore behavior are required for strong production reliability.
The Wiki explains infrastructure responsibilities, topology, ownership, and ecosystem boundaries.
The victus-infra repository defines exact Compose services, Ansible behavior, host paths, firewall rules, DNS records, secret names, deployment workflows, and operational procedures.