AI-Chain

Kubernetes Is More Than a Deployment Tool: Turning Containers into a Self-Healing Platform with Control Loops

Share:
Kubernetes Is More Than a Deployment Tool: Turning Containers into a Self-Healing Platform with Control Loops
# Kubernetes Is More Than a Deployment Tool: Turning Containers into a Self-Healing Platform with Control Loops Many teams first meet Kubernetes through a list of commands: create a Deployment, expose a Service, set the replica count, and hope the application keeps running. Kubernetes is worth studying for a deeper reason. It turns day-to-day distributed-systems operations into continuously running control loops. You describe the state you want. Kubernetes observes the current state, calculates the difference, and uses controllers, the Scheduler, kubelet, and networking components to move the system toward that target. If a container exits, it can be restarted. If replicas are missing, they can be recreated. During a version update, a Deployment can replace old replicas gradually. The model changes deployment from a one-time action into continuous maintenance. This article uses `kubernetes/kubernetes` as its subject. It starts with the repository and then explains the relationship between the control plane and workloads. A reproducible local example covers deployment, service discovery, health checks, configuration, rolling updates, and rollback. The goal is not to memorize YAML, but to understand the role of every resource in the control loop. ## Why Kubernetes Is Still Worth Studying As of August 10, 2026, the GitHub API reported about 124,378 stars for `kubernetes/kubernetes`. Its latest push was on August 9, 2026, and the repository uses the Apache-2.0 license. The same API response listed `v1.37.0-rc.0`, published on August 6, 2026. These facts confirm project activity and release context; they do not mean that a release candidate is appropriate for every production environment. The official README describes Kubernetes as an open source system for managing containerized applications across multiple hosts, with mechanisms for deployment, maintenance, and scaling. The official concepts documentation adds service discovery and load balancing, self-healing, horizontal scaling, rolling updates, and configuration management. It is a strong implementation-focused project for three reasons: 1. **It is an executable control platform.** It includes the API Server, Controller Manager, Scheduler, kubelet, and a large set of controllers rather than being only a specification or resource catalog. 2. **Its abstractions are observable.** You can inspect resources with `kubectl get`, trace events with `describe`, and compare a YAML declaration with the resulting state. 3. **The same model extends to AI services.** APIs, workloads, services, resource limits, and health checks are also the basic engineering problems of model-serving, data-processing, and Agent backends. ## The Core Is the Control Loop, Not the Pod Beginners often equate Kubernetes with a tool for starting Pods, but a Pod is only the smallest deployable unit. In practice, an application is described by several objects. A Deployment manages a stateless workload, a ReplicaSet maintains the replica count, a Service provides a stable network entry point, ConfigMaps and Secrets separate configuration from sensitive data, and probes give the platform a way to judge application health. A deployment can be understood as this sequence: 1. The user sends a desired state to the API Server, such as “run three `web` Pods with image version `v2`.” 2. The Deployment Controller creates or updates a ReplicaSet. 3. The ReplicaSet Controller observes the number of Pods and creates more when the count is below three. 4. The Scheduler selects a suitable Node for an unbound Pod using resource requests, limits, affinity, and other conditions. 5. The kubelet on the Node asks the container runtime to start the container and keeps reporting status. 6. If the liveness or readiness conditions fail, the platform can restart a container or stop sending traffic to it, depending on the probe. The key idea is that a controller is not a script. A script runs and exits; a controller continues watching resources and events. If a Pod is deleted, the system can observe the difference and try to repair it. That is why Kubernetes can handle node failures, interrupted rollouts, and replica drift. ## Build a Small Observable Example The following example uses Minikube as a local learning environment. The official Hello Minikube tutorial covers cluster creation, Deployment, Service exposure, inspecting Pods and Nodes, scaling, and rolling updates. If you already have another Kubernetes cluster, replace the Minikube commands with the equivalent commands for that cluster. ### Check the Environment First Prepare Docker or another compatible container runtime, together with `kubectl` and Minikube. Start the cluster and check the Node: ```bash minikube start kubectl cluster-info kubectl get nodes -o wide ``` If `minikube start` fails, first check whether the container runtime is running, whether the selected driver is available, and whether the machine has enough CPU, memory, and disk. Do not blame the YAML before confirming that the cluster is Ready; every later layer depends on it. ### Describe the Workload with a Deployment Create `web.yaml`: ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: web labels: app: web spec: replicas: 2 selector: matchLabels: app: web template: metadata: labels: app: web spec: containers: - name: nginx image: nginx:stable-alpine ports: - name: http containerPort: 80 resources: requests: cpu: 100m memory: 128Mi limits: cpu: 500m memory: 256Mi readinessProbe: httpGet: path: / port: http periodSeconds: 5 livenessProbe: httpGet: path: / port: http initialDelaySeconds: 10 periodSeconds: 10 ``` Apply it and observe the result: ```bash kubectl apply -f web.yaml kubectl get deployment,replicaset,pod -l app=web kubectl describe deployment web kubectl get events --sort-by=.lastTimestamp ``` These commands expose different levels of the system. `get` shows a summary, `describe` shows the controller's computed status and events, and `events` helps identify scheduling, image-pull, or probe failures. For a `Pending` Pod, check resources and scheduling. For `ImagePullBackOff`, check the image name, tag, and registry. For `CrashLoopBackOff`, inspect container logs and startup arguments. ### Give the Workload a Stable Network Entry Point Pods can be recreated and their IP addresses can change, so other services should not depend on Pod IPs directly. Create `web-service.yaml`: ```yaml apiVersion: v1 kind: Service metadata: name: web spec: selector: app: web ports: - name: http port: 80 targetPort: http type: ClusterIP ``` Run: ```bash kubectl apply -f web-service.yaml kubectl get service web kubectl get endpointslices -l kubernetes.io/service-name=web minikube service web --url ``` A Service finds matching Pods through its selector and provides a stable service abstraction. This distinction matters: the Deployment manages which replicas exist, while the Service manages where traffic goes. Mixing those responsibilities makes debugging much harder. ## Use Three Kinds of Health Check for Different Problems The official documentation separates probes into liveness, readiness, and startup probes: - **livenessProbe** asks whether the application has entered a state it cannot recover from. A failed liveness probe can cause kubelet to restart the container. - **readinessProbe** asks whether the application can accept traffic right now. A failed readiness probe removes the Pod from available Service endpoints without necessarily restarting it. - **startupProbe** asks whether a slow-starting application has finished starting. Until it succeeds, liveness and readiness checks are not activated, which is useful for long warm-up periods. This separation prevents a common mistake in AI inference services: treating “the model is still loading” as “the process is broken.” Model loading, GPU initialization, and cache creation can take a long time. The process may be alive but not ready for traffic. Combining startup and readiness probes is usually clearer than setting an extremely large liveness delay. A probe should also be cheap and predictable. An HTTP health endpoint should not scan an entire database, run inference, or call an expensive external API every time. Otherwise, the platform may create load while trying to decide whether the service is healthy. ## Separate Configuration, Secrets, and Image Versions Hard-coding all configuration into an image makes the same image difficult to reuse across environments. ConfigMaps are suitable for non-sensitive values such as ports, feature flags, and environment names. Secrets are used to pass tokens, certificates, and other sensitive values. A Secret is not a complete secret-management system; production environments still need to consider encryption at rest, RBAC, an external secret manager, rotation, and auditing. Here is an example containing no real credential: ```yaml apiVersion: v1 kind: ConfigMap metadata: name: web-config data: LOG_LEVEL: info FEATURE_MODE: production --- apiVersion: v1 kind: Secret metadata: name: web-secret type: Opaque stringData: API_TOKEN: "[REDACTED]" ``` After applying it, inject values with `envFrom` or individual `env` entries. Be careful while inspecting configuration: `kubectl describe`, events, shell history, and CI logs can all expose sensitive data. Do not paste `kubectl get secret -o yaml` into an issue or chat channel. ## Rolling Updates, Scaling, and Rollback Kubernetes becomes more valuable after the first deployment, when the system must keep operating during change. Update the image: ```bash kubectl set image deployment/web nginx=nginx:1.27-alpine kubectl rollout status deployment/web kubectl rollout history deployment/web ``` If the new version behaves incorrectly, roll it back: ```bash kubectl rollout undo deployment/web kubectl rollout status deployment/web ``` Behind these commands, the Deployment Controller creates a new ReplicaSet and gradually adjusts the old and new replica counts. Availability also depends on `maxUnavailable`, `maxSurge`, PodDisruptionBudget, probes, and whether the application shuts down gracefully. A completed rollout is not the same as a completed feature validation; combine it with smoke tests, metrics, logs, and error-rate checks. Horizontal scaling changes the desired replica count: ```bash kubectl scale deployment/web --replicas=4 kubectl get pods -l app=web -w ``` Automatic scaling needs the appropriate metrics components and configuration. Creating an HPA object alone does not guarantee that it can make a decision. For model-serving workloads, GPU memory, concurrent requests, batch size, and model load time may describe capacity better than CPU utilization, so the scaling policy must match the service. ## Diagnose the Control Loop from Failure Symptoms A fixed troubleshooting order is more useful than memorizing more commands. First check whether the resources exist: `kubectl get deployment web`, `kubectl get pods -l app=web`, and `kubectl get svc web` confirm that the API Server accepted the objects. Next check convergence: the Deployment's `AVAILABLE` count, the Pods' `READY` state, and the Service's EndpointSlices should agree. Then inspect events and logs: `kubectl describe pod` shows scheduling, mounts, probes, and image-pull events, while `kubectl logs` identifies application startup errors. The symptoms can be separated into layers. A `Pending` Pod usually leads to Node resources, taints, affinity, or PVC checks. `ContainerCreating` points toward images, volumes, CNI, or Secrets. Repeated restarts require the previous container's logs and exit code. A new ReplicaSet without available replicas suggests checking readiness probes, resource limits, and image versions. A Service without endpoints suggests comparing its selector with Pod labels exactly. The purpose is to turn “the service is down” into API, scheduling, startup, health, and networking hypotheses that can be verified one by one. After a change, perform both positive and negative verification. Confirm that the healthy version is reachable, then deliberately deploy a nonexistent image or temporarily increase the replica count, observe events and rollout status, restore the configuration, and confirm that the system returns to a stable state. That proves more than seeing a single “deployment successfully configured” message. ## Production Boundaries That Are Easy to Miss ### 1. Resource Requests Are Not Decoration Requests affect scheduling, and limits affect the resources available to a container. Without reasonable requests, the Scheduler cannot make reliable decisions. Without usage measurements, limits can cause OOMKills or CPU throttling. Establish a baseline with metrics before tuning; do not copy someone else's numbers. ### 2. Namespaces and RBAC Belong in the Same Design Namespaces help separate teams, environments, and quotas, but they are not a complete security boundary. ServiceAccounts, Roles, RoleBindings, and least privilege control API access. A workload should not receive cluster-wide read-write permissions by default. ### 3. Observability Should Follow the Control Loop In addition to application logs, inspect Pod status, available Deployment replicas, scheduling events, probe failures, node pressure, and API Server errors. When something fails, ask what the desired state is, what the current state is, and which controller has not converged the difference. This is usually faster than restarting everything blindly. ### 4. YAML Maintainability Determines Change Speed A single file is fine for a small experiment. Team operations need version control, environment overlays, validation, and review. Whether you use Kustomize, Helm, or another tool, the final resources sent to the API Server should be traceable, reproducible, and rollback-friendly. Templates are not the goal; the goal is to answer which workloads a change will affect. ## An AI Chain Perspective: Treat Model Serving as Part of the Platform Kubernetes does not solve model quality, prompt design, or data governance. It can, however, carry the engineering problems that appear when an AI application enters production: - A **model API** can use a Deployment for replicas and a Service for a stable endpoint. - **Model loading time** can be represented with a startup probe so that an unready container is not restarted or sent traffic prematurely. - **Version changes** can use multiple Deployments, labels, and traffic strategies instead of replacing every instance at once. - **GPU workloads** can use Kubernetes resource and scheduling capabilities, together with the correct device plugin, node labels, resource requests, and monitoring. - **Data-processing tasks** can use Jobs or CronJobs to describe one-off and scheduled work, with retry, retention, and output behavior made explicit. This decomposition lets an AI system grow from “one service on one machine” into an observable, updateable, recoverable engineering system. Cluster complexity is still a real cost. For a small service, Docker Compose or a managed platform may be a better fit. Choose Kubernetes because you need its control loops, scaling behavior, or ecosystem—not only because it is popular. ## Conclusion: Learn to Observe the Difference Before Memorizing YAML The most valuable Kubernetes lesson is not a particular `kubectl` flag. It is the engineering idea of declaring a desired state and letting control loops converge on it. Deployments manage replicas and versions, Services manage stable endpoints, probes manage health decisions, and ConfigMaps and Secrets define configuration boundaries. Each object should have a clear responsibility and a way to be verified through commands, events, and metrics. Start with the small example in this article. Create two replicas, confirm that the Service finds their endpoints, deliberately break the image and observe the rollout, then roll back and inspect the events. When you can explain what difference a controller saw, what it did, and how to verify the next state, you are no longer merely applying YAML—you are beginning to understand Kubernetes. ## Sources - Kubernetes GitHub repository: - Kubernetes README: - Kubernetes Overview: - Deployments: - Service: - Secrets: - Liveness, Readiness, and Startup Probes: - kubectl Quick Reference: - Hello Minikube: