Replicator
Replicator is a container you run in your own infrastructure. It holds a connection to Schematic’s datastream, loads all of your flag, company, and user data into a Redis instance you host, and keeps that copy current in real time. Backend SDKs in Replicator mode then evaluate flags from Redis without calling the Schematic API, so flag checks stay fast and keep working if Schematic is unreachable.
Setting it up takes three pieces:
- A Redis instance (single node or cluster) that both the Replicator and your application can reach.
- The Replicator container, configured with your Schematic API key and Redis address.
- Your backend SDK, configured for Replicator mode and pointed at the same Redis. See your SDK’s page: Go, Node.js, Python, Java, C#, or Ruby.
Image
The Replicator is published as a multi-architecture image (linux/amd64 and linux/arm64) to two registries:
Each release is tagged with its full version (0.2.7), its minor version (0.2), its major version (0), and latest. In production, pin a full version (0.2.7) so a new release only reaches you when you choose to upgrade. A minor version tag (0.2, as in the examples below) moves with each patch release, so it picks up fixes automatically the next time the image is pulled.
Quick start
The Replicator exits at startup if it cannot connect to Redis, so make sure Redis is reachable from the container first. It also needs outbound HTTPS access to two hosts:
api.schematichq.com, for the initial load of companies and users.datastream.schematichq.com, for the datastream WebSocket (wss://).
Use a secret (backend) API key for SCHEMATIC_API_KEY, and the key for the environment whose data you want replicated. Store it in your platform’s secret manager rather than in plain configuration.
Deployment
One writer per Redis
Exactly one Replicator instance may write to a given Redis. A second writer would double-write the cache and corrupt the cursor the Replicator uses to resume after a reconnect.
The Replicator enforces this with a lease in Redis. An instance acquires the lease at startup, renews it while it runs, and releases it on shutdown. A crashed instance’s lease expires after WRITER_LOCK_TTL (15 seconds by default) so a replacement can take over.
A starting instance that finds the lease held waits up to WRITER_LOCK_ACQUIRE_TIMEOUT (5 minutes by default) for it to be released, then exits with an error. While it waits, /health returns 200 and /ready returns 503. In practice:
- Run one replica per Redis.
- Use a deployment strategy that stops the old instance before or while the new one waits. On Kubernetes, use
strategy: Recreate: the defaultRollingUpdatewaits for the new pod to become ready before stopping the old one, which never happens while the old pod holds the lease. - On platforms that drain and stop the old task once the new one passes a liveness check (for example, Amazon ECS with a load balancer health check on
/health), a rolling deploy hands the lease over within the acquire timeout. - For high availability, run Redis in a highly available configuration and let your orchestrator restart a failed Replicator promptly. Do not run several Replicators against the same Redis.
- To run more than one Replicator, for example one per Schematic environment, give each its own Redis, or its own
REDIS_DBon a single-node Redis, and point each environment’s SDKs at the matching one. ChangingWRITER_LOCK_KEYdoes not separate them: every Replicator writes the sameschematic:keys, so two Replicators in one Redis database overwrite each other’s data.
Docker Compose
Runs the Replicator alongside a Redis instance:
Kubernetes
Add resource requests and limits based on the usage you observe. Memory use grows with the number of companies and users, and with the buffer sizes in Performance tuning.
Health checks
In the default asynchronous loading mode, /ready returns 200 as soon as the datastream subscription is set up, while companies and users are still loading in the background (see Initial load and reconnects). With USE_ASYNC_LOADING=false, it waits for the full load.
Both endpoints return the same JSON shape:
components.datastream on /ready is ready, connected_loading (connected, still setting up), or disconnected; on /health it is connected, disconnected, or unknown (still starting). status is always healthy and components.redis is always connected; use the status code, not these fields, to judge health. cache_version is an 8-character hash that forms the version segment of the Redis keys the Replicator writes (see Redis key layout).
The image has no HTTP client installed. For a container-level health check, use the binary’s built-in probe, which exits 0 when the endpoint returns 200 and respects HEALTH_PORT:
Configuration
The Replicator is configured entirely through environment variables. Durations use Go’s format: 500ms, 45s, 30m, 1h30m.
Core
Redis
Writer lease
Initial load and reconnects
By default the Replicator loads all companies and users from the Schematic API in the background, so the datastream connects quickly. It saves its position in the datastream to Redis, so after a reconnect or a restart it resumes from the last message it processed instead of reloading everything. A full reload happens only on a first start against an empty Redis, or when the saved position is too old to resume from.
Connection keepalive
The Replicator pings the datastream to keep the WebSocket open through load balancers and proxies with idle timeouts.
The defaults suit an idle timeout of 60 seconds, which is common for cloud load balancers. For a 30-second idle timeout, use WS_PING_INTERVAL=15s and WS_PONG_WAIT=25s. For 120 seconds or more, 50s and 60s reduce traffic.
Performance tuning
These control how updates from the datastream are batched and written to Redis. The defaults favor low latency and suit most deployments.
Starting points for different environments:
Smaller batches and shorter timeouts lower latency; larger ones raise throughput. Reduce channel sizes if memory is constrained; total buffering is roughly NUM_WORKERS × (COMPANY_CHANNEL_SIZE + USER_CHANNEL_SIZE).
While the circuit breaker is open, updates arriving from the datastream are dropped, not queued, so Redis falls behind. They are not lost for good: the Replicator’s saved position never moves past an update it failed to apply, so the next reconnect or restart replays them, and if too many build up it starts a full reload on its own. If you see Circuit breaker open, dropping in the logs, fix the Redis problem, then restart the Replicator to catch up immediately.
Tracing
The Replicator can export traces over OTLP to any compatible collector or backend, such as the OpenTelemetry Collector, Datadog, Jaeger, or Honeycomb. Tracing is off until an endpoint is set, and it uses the standard OpenTelemetry environment variables:
The Replicator emits one trace per batch of updates: a root span (replicate company, replicate user, or replicate flags) with a child span for each Redis write or delete. Every span carries replicator.entity and replicator.batch.size. A batch dropped because the circuit breaker is open is recorded as a failed span with replicator.circuit_breaker.open. All spans are sampled; configure sampling in your collector.
Redis key layout
The Replicator writes, and SDKs in Replicator mode read, these keys. <version> is the cache_version reported by /health.
The schematic: prefix is fixed. If your SDK’s Redis cache is configured with a different prefix, it will never find these keys and will silently fall back to calling the Schematic API. Leave the SDK’s prefix at its default.
Troubleshooting
The container exits immediately. Check the logs with docker logs schematic-replicator. The two most common causes are a missing SCHEMATIC_API_KEY and an unreachable Redis.
The container exits after a few minutes with “Could not acquire writer lock”. Another Replicator is holding the lease on this Redis. Make sure only one instance runs, and on Kubernetes use strategy: Recreate. See One writer per Redis.
/ready stays at 503. Check the logs first: the Replicator may be waiting for the writer lease (see above). Otherwise components.datastream is usually disconnected, which means the API key is wrong or the container cannot reach datastream.schematichq.com over wss://.
The connection drops every 30 or 60 seconds. A load balancer or proxy is closing idle connections. Lower WS_PING_INTERVAL and WS_PONG_WAIT; see Connection keepalive.
Your SDK still calls the Schematic API. Confirm the SDK is in Replicator mode, can reach the Replicator’s /ready endpoint, uses the same Redis, and uses the default key prefix. Set LOG_LEVEL=debug on the Replicator to log each batch it writes to Redis.