← All articles
Information TechnologyMay 6, 20238 min read

Kubernetes Architecture and its Components

How the control plane, worker nodes, and cluster services fit together to orchestrate containers at scale.

The big picture

A Kubernetes cluster has two halves: the control plane, which makes decisions, and the worker nodes, which run your workloads. The control plane never runs your apps; the nodes never make global decisions. Between them sits a single source of truth — etcd — and a single front door — the API server. Everything else is a client of those two.

etcd: the source of truth

etcd is a distributed key-value store holding every object in the cluster: Pods, Deployments, Secrets, the lot. It uses the Raft consensus protocol, which means it survives the loss of a minority of members — run 3 or 5 members, never 2 or 4 (even numbers don't buy you extra fault tolerance, just extra latency).

  • Quorum math: a 3-member cluster tolerates 1 failure; 5 members tolerate 2. Lose quorum and the cluster goes read-only — nothing can be scheduled or changed.
  • Back it up: etcdctl snapshot save on a schedule, stored off-cluster. etcd holds your Secrets; an untested backup is not a backup.
  • Encrypt Secrets at rest: etcd stores them base64-encoded by default — readable to anyone with etcd access. Enable encryption providers.

kube-apiserver: the only one who talks to etcd

The API server is the only component allowed to touch etcd directly. Every create, read, update, and delete flows through it, which makes it the cluster's choke point and its policy engine:

  • Authentication: who are you? (certs, tokens, OIDC)
  • Authorization: what may you do? (RBAC — Roles and Bindings)
  • Admission control: mutating and validating webhooks — where policy engines like Kyverno or OPA Gatekeeper plug in.

It is stateless and horizontally scalable — run several behind a load balancer. If the API server is down, running workloads keep running, but nothing can change: no deploys, no scaling, no self-healing.

kube-scheduler: matchmaking for Pods

The scheduler watches for unscheduled Pods and assigns each to a node in two phases: filtering (which nodes can run it — enough CPU/memory? matching taints/tolerations? node selectors?) then scoring (which candidate is best — least-loaded wins by default). You steer it with:

  • nodeAffinity: prefer (or require) nodes with certain labels — e.g. SSD nodes, or a specific zone.
  • podAntiAffinity: spread replicas across nodes or zones — the single most important setting for real high availability.
  • taints and tolerations: keep general workloads off special nodes (GPU nodes, infra nodes).
If all three replicas of your app land on the same node, you don't have high availability — you have a deployment-shaped single point of failure. podAntiAffinity is the fix.

kube-controller-manager: the control loops

This is a bundle of controllers, each running the same tiny loop: observe actual state → compare with desired state → act to close the gap. The Node controller evicts Pods from dead nodes; the Replication controller keeps replica counts; the Deployment controller orchestrates rollouts. Kubernetes' famous "self-healing" is just dozens of these dumb, relentless loops.

kubelet: the node agent

On every worker node, the kubelet is the local enforcer. It watches the API server for Pods assigned to its node and makes them real: pulling images through the container runtime (containerd, via the CRI), mounting volumes, executing health probes, and reporting status back. If a liveness probe fails, the kubelet restarts the container. If the node stops heartbeating, the control plane notices — and the controllers start rescheduling elsewhere.

kube-proxy: the cluster's load balancer

kube-proxy implements Service networking on each node — programming iptables (or IPVS at scale) so that a Service's virtual IP load-balances across healthy Pod endpoints. Modern clusters often replace it with eBPF-based networking (Cilium), which does the same job faster and with richer policy. Either way: the Service IP never belongs to a real interface; it's a fiction maintained identically on every node.

The plugin model: CNI, CRI, CSI

Kubernetes stays small by delegating through three plugin interfaces:

  • CRI (runtime): containerd, CRI-O — actually runs containers.
  • CNI (networking): Calico, Cilium, Flannel — gives Pods IPs and enforces NetworkPolicies.
  • CSI (storage): EBS, Azure Disk, Ceph drivers — provisions PersistentVolumes on demand.

Your choice of CNI especially shapes what the cluster can do — Cilium's eBPF datapath, for instance, unlocks observability and security features kube-proxy never had.

Cluster-wide services

  • DNS (CoreDNS): every Service gets a DNS name automatically — service discovery with zero config.
  • Metrics Server: feeds CPU/memory metrics to kubectl top and the Horizontal Pod Autoscaler.
  • Ingress controller + cert-manager: external HTTP routing and automatic TLS.

High availability: what "production" actually means

A real cluster runs 3+ control-plane nodes (stacked etcd, or external etcd for large clusters) behind a load balancer, spread across failure domains. Worker nodes run in multiple availability zones with podAntiAffinity spreading critical workloads. The failure drills worth running:

  • Lose a worker node: Pods reschedule after the node-monitor grace period (~40s default). PDBs (PodDisruptionBudgets) protect voluntary disruptions, not this.
  • Lose etcd quorum: the cluster freezes — workloads run, but nothing changes until quorum returns. This is why etcd backups and odd member counts matter.
  • Lose the API server: same freeze. Run it replicated; it's stateless.

Putting it together

kubectl apply → API server validates and writes to etcd → controllers notice the new desired state → scheduler assigns Pods to nodes → kubelets pull images and start containers → kube-proxy wires up Service IPs → CoreDNS makes them resolvable. Every component does one job, watches the shared state, and retries forever. Simple parts, composed relentlessly — that's the architecture.

MK
Mohankrishna PodileDevOps Engineer & Cloud Architect · Irving, Texas

Comments

Questions, corrections, war stories — all welcome. Sign in with GitHub to join the discussion.