Skip to content

Factories > Managed self-hosting

Self-hosting troubleshooting

Open in ChatGPT ↗
Ask ChatGPT about this page
Open in Claude ↗
Ask Claude about this page
Copied!

Diagnose and fix common problems with self-hosted Automation Platform worker daemons across Docker, Kubernetes, and Direct backends.

Use these checks when the oz-agent-worker daemon won’t start or connect, tasks stay queued, or tasks fail.


Cause: Docker isn’t running, or the daemon platform isn’t supported.

Fix:

  1. Verify Docker is running: docker info.
  2. Confirm the daemon platform is linux/amd64 or linux/arm64. Windows containers are not supported.
  3. If the worker runs inside Docker with a rootful Linux daemon, mount /var/run/docker.sock and pass its numeric group ID with --group-add. See the Docker installation example.

Cause: The worker Deployment couldn’t start, reach the Kubernetes API, or create its preflight Job.

Fix:

  1. Run kubectl describe pod -n NAMESPACE WORKER_POD. Replace NAMESPACE with the chart namespace and WORKER_POD with the worker pod name. For CreateContainerConfigError, verify the Secret and key configured by warp.apiKeySecret.
  2. Check the worker logs for Kubernetes API or preflight diagnostics: kubectl logs -n NAMESPACE WORKER_POD.
  3. Confirm the worker’s namespace has these permissions: create, get, list, watch, delete on jobs; get, list, watch on pods; get on pods/log; list on events.
  4. Confirm the task namespace allows pods with a root init container, unless you enabled native image volumes with kubernetesBackend.useImageVolumes=true.
  5. If your cluster restricts image sources, set kubernetesBackend.preflightImage to an allowlisted image. The default is busybox:1.36.
  6. To pull the preflight image from a private registry, configure imagePullSecrets in kubernetesBackend.podTemplate.

A successful preflight validates the configured pod shape, not task-specific images, Secrets, setup commands, or network access.

Cause: The oz CLI isn’t installed or isn’t on the worker’s PATH.

Fix:

  1. Install the Oz CLI on the worker host. See Installing the CLI.
  2. If the CLI isn’t on PATH, set oz_path in the config file to the absolute path of the oz binary.

Cause: The API key is invalid, expired, or the host cannot reach the Automation Platform‘s backend.

Fix:

  1. Confirm you created a Self-hosted worker API key and that it has not expired.
  2. If you suspect the key is invalid, open the Warp Factories web app user settings page, click Generate new token, and select Self-hosted worker to create a replacement.
  3. Ensure the host has outbound internet access to oz.warp.dev:443.
  4. Check that no firewall rules are blocking WebSocket connections to wss://oz.warp.dev.
  5. Increase log verbosity with --log-level debug to see connection details.

See Security and networking for the full list of outbound endpoints the worker needs.


Cause: The worker isn’t running, the --host value doesn’t match the worker’s --worker-id, or the worker and task belong to different teams.

Fix:

  1. Confirm the worker is running and connected. Check the worker logs for Successfully connected to server.
  2. Verify the --host (or worker_host) value you passed matches your --worker-id exactly. Case-sensitive.
  3. Ensure the worker’s team matches the team creating the task.

Cause: The worker is running, but its metrics don’t reach Prometheus or your collector.

Fix:

  1. Confirm OTEL_METRICS_EXPORTER is set on the worker process. In prometheus mode, run curl -s localhost:9464/metrics from the worker host to check that the endpoint responds.
  2. For Prometheus scrape mode, confirm the bind address is 0.0.0.0 (not localhost) when running in Docker or Kubernetes. localhost is only reachable from inside the container.
  3. Confirm no firewall or network policy blocks the metrics port (9464 by default).
  4. In OTLP push mode, confirm OTEL_EXPORTER_OTLP_ENDPOINT points to a reachable collector and the protocol matches (http/protobuf or grpc).
  5. With the Helm chart, confirm metrics.enabled=true, then check that the Service and any PodMonitor exist: kubectl get svc,podmonitor -n NAMESPACE. Replace NAMESPACE with the chart namespace. A PodMonitor needs the Prometheus Operator CRDs (monitoring.coreos.com) installed in the cluster.
  6. Restart the worker with --log-level debug and look for metrics errors at startup.

See Monitoring for the full setup guide.


Cause: The task’s environment, resources, or dependencies failed. The logs show which.

Fix (all backends):

  1. Review task logs in the cloud agent dashboard or through session sharing.
  2. To keep the container, Job, or workspace for inspection, run the worker with --no-cleanup. On Kubernetes with cleanup enabled, failed Jobs stay for 24 hours by default.
  3. Run the worker with --log-level debug for detailed execution logs.
  1. Verify Docker is running (docker info).
  2. If using a custom image, confirm it is glibc-based (not Alpine/musl) and that its architecture matches the worker’s Docker daemon platform.

The Helm chart’s worker.resources sizes the worker Deployment, not task Jobs. To size task containers, see Size task containers.

Start with the state of the Job and its Pod:

Terminal window
kubectl get jobs,pods -n NAMESPACE
kubectl describe pod -n NAMESPACE TASK_POD
kubectl logs -n NAMESPACE TASK_POD -c CONTAINER_NAME
kubectl logs -n NAMESPACE TASK_POD -c CONTAINER_NAME --previous

Replace NAMESPACE with the task namespace, TASK_POD with the task pod name, and CONTAINER_NAME with the name of the container that failed. Then find the symptom below.

No node has room for the Pod. The worker fails the task if the Pod stays unschedulable longer than kubernetesBackend.unschedulableTimeout, which defaults to 10 minutes.

Verify: In kubectl describe pod, look for PodScheduled=False, FailedScheduling, and Insufficient cpu or Insufficient memory. Compare the Pod’s requests with free node capacity, then check selectors, affinity, taints and tolerations, topology constraints, quotas, and volume binding.

Fix:

  1. Lower the task’s CPU and memory requests, lower worker.maxConcurrentTasks, or add node capacity.
  2. If you use a node autoscaler, make sure it can add nodes that match the Pod’s selectors, affinity, and tolerations. If provisioning takes longer than 10 minutes, raise unschedulableTimeout.

Raising only a limit doesn’t help, because the scheduler places Pods by their requests. An instance shape sets the request equal to the limit, so a larger shape needs a larger free node.

The task container used more memory than its limit.

Verify: In kubectl describe pod, confirm the task container’s last state shows reason OOMKilled. Compare the workload’s peak memory with the container’s memory limit.

Fix: Reduce the workload’s peak memory, or raise the task’s memory in the runner’s instance shape or in the task container of pod_template. Init containers have their own resources and are sized separately.

The node evicted the Pod, usually because of memory, disk, or other node pressure.

Verify: Read the Pod’s reason and events for the pressure type.

Fix: Free up node resources, lower concurrency, or add capacity, then rerun the task. A replacement Pod can’t recover the task’s emptyDir workspace. If voluntary disruption caused the eviction, see Protect active task pods from disruption.

Exit code 143 means the process received SIGTERM. It doesn’t indicate an out-of-memory failure. Kubernetes reports those as OOMKilled.

Verify: Check the container’s termination reason and the Pod’s events for eviction, preemption, a node drain, manual deletion, or the Job deadline. kubernetesBackend.activeDeadlineSeconds defaults to eight hours, after which Kubernetes terminates the Job with DeadlineExceeded.

Fix: Address the cause the events show. If the Job hit its deadline, raise activeDeadlineSeconds or split the task.

  • ErrImagePull, ImagePullBackOff, or InvalidImageName - See Image pull failures. Preflight doesn’t pull task images.
  • CreateContainerConfigError or FailedMount - The Pod’s events name the missing Secret, ConfigMap, service account, key, or volume. Check that it exists in the task namespace.
  • Init container failure - Check each init container’s status and logs. The init container that loads the Warp sidecar runs as root unless you enable native image volumes. Custom init containers must finish before the task starts.
  • Network failure - Test DNS, TLS, and the destination from the task Pod, not the worker Pod. A working worker connection to Warp says nothing about the task’s network policies, service mesh, proxy, or egress.
  • Missing task credentials - Provide repository, registry, and application credentials through your Secret integration and pod_template. The worker API key authenticates the worker to Warp and isn’t available to tasks.
  1. Verify the Oz CLI is accessible.
  2. Verify the workspace root directory has write permissions for the user running the worker.

  1. If using a private registry, ensure Docker credentials are available to the worker. See Private Docker registries.
  2. Try pulling the image manually on the worker host: docker pull <image>.
  1. Configure imagePullSecrets in the pod_template section of your worker config.
  2. Verify the Secret exists in the task namespace and contains valid credentials.
  • Verify the image exists and the tag is correct.
  • Check network connectivity from the worker/cluster to the registry.