Post

K8s CNI(1) β€” Cilium vs OVN and DPU Offloading

K8s CNI(1) β€” Cilium vs OVN and DPU Offloading

πŸ” Are Cilium and OVN Competing Kubernetes Network Solutions?

Broadly, yes. Cilium implements Kubernetes networking, security, and observability through eBPF programs in the Linux kernel. OVN-Kubernetes represents the cluster as a logical network and translates that model into OVS flows. Cilium is generally well suited to cloud-native Kubernetes environments, while OVN has a particularly mature path for virtualization, OpenShift, and DPU hardware offloading.


β‘  CNI Is the Interface, Not the Dataplane

Kubernetes requires a network implementation to connect Pods, assign IP addresses, and configure routes. It requests these operations through the Container Network Interface, or CNI.

1
2
3
4
5
6
7
kubelet
    ↓
CNI Request
    ↓
Network Implementation
    ↓
Pod Interface and Route

CNI defines how Kubernetes requests network configuration. It does not define how packets must be forwarded after that configuration is complete. Different CNI implementations can therefore use different dataplanes.

1
2
3
4
5
6
7
8
9
10
Kubernetes CNI
    β”‚
    β”œβ”€β”€ Calico
    β”‚     └── Linux routing / iptables / eBPF
    β”‚
    β”œβ”€β”€ Cilium
    β”‚     └── eBPF
    β”‚
    └── OVN-Kubernetes
          └── OVN logical network / OVS flows

Cilium and OVN-Kubernetes solve many of the same Kubernetes networking problems, but they approach them from different architectural directions.


β‘‘ How Cilium Works

Cilium is a networking, security, and observability platform built around eBPF. eBPF allows verified programs to run at predefined points inside the Linux kernel. In networking, these points include TC, XDP, sockets, and cgroups. A simplified Cilium path looks like this.

1
2
3
4
5
6
7
8
9
10
11
Kubernetes API
        ↓
cilium-operator
        ↓
cilium-agent on each node
        ↓
eBPF Programs and Maps
        ↓
Linux Kernel
        ↓
NIC

The cilium-agent watches Kubernetes resources and programs the local node’s eBPF datapath. For example, it can install logic for:

  • Pod forwarding,
  • NetworkPolicy enforcement,
  • Service load balancing,
  • source NAT,
  • encryption,
  • traffic observability.

Packets are processed by eBPF programs attached to the relevant kernel hooks.

1
2
3
4
5
6
7
8
9
10
11
Pod
 ↓
veth
 ↓
TC eBPF
 β”œβ”€β”€ Policy
 β”œβ”€β”€ Routing
 β”œβ”€β”€ Service Load Balancing
 └── Observability
 ↓
NIC

Cilium can also replace kube-proxy. A conventional Kubernetes Service path may use iptables or IPVS rules maintained by kube-proxy.

1
2
3
4
5
6
7
Client
  ↓
Service IP
  ↓
kube-proxy rules
  ↓
Backend Pod

With Cilium’s kube-proxy replacement, Service translation is implemented through eBPF maps and programs.

1
2
3
4
5
6
7
Client
  ↓
eBPF Service Lookup
  ↓
Backend Selection
  ↓
Backend Pod

Cilium’s major operational advantage is that networking, policy, Service handling, and observability can be implemented within one integrated dataplane.


β‘’ How OVN-Kubernetes Works

OVN-Kubernetes uses OVN to model the cluster as a logical network and OVS to execute the resulting forwarding rules. Its architecture has more explicit control-plane layers than Cilium.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
Kubernetes API
        ↓
OVN-Kubernetes Controller
        ↓
OVN Northbound Database
        ↓
ovn-northd
        ↓
OVN Southbound Database
        ↓
ovn-controller
        ↓
Open vSwitch
        ↓
NIC

The Northbound database stores the intended logical network.

Examples include:

  • logical switches
  • logical routers
  • ACLs
  • load balancers
  • logical ports
1
2
3
4
5
6
Desired Network

Logical Switch A
Logical Router R
ACL: allow namespace-a
Load Balancer: service-x

ovn-northd does not simply copy this information. It compiles the logical network into lower-level logical flows and bindings stored in the Southbound database.

1
2
3
4
5
Northbound Intent
        ↓
ovn-northd
        ↓
Southbound Logical Flows

Each node runs ovn-controller, which reads the relevant Southbound state and programs OVS.

1
2
3
4
5
6
7
Southbound Database
        ↓
ovn-controller
        ↓
OVS Flow Rules
        ↓
Packet Forwarding

OVN therefore separates the desired logical topology from the node-local forwarding implementation.


β‘£ Cilium vs OVN

The most important difference is not simply eBPF versus OVS. It is the abstraction each system treats as its center.

CiliumOVN-Kubernetes
Centers on programmable Linux kernel hooksCenters on a logical network model
Programs eBPF maps and programs per nodeCompiles logical switches and routers into OVS flows
Kubernetes-native architectureOriginates from virtualized SDN architecture
Strong integration of networking, policy, Service, and observabilityStrong logical topology, virtual routing, and network isolation model
Commonly uses TC, XDP, sockets, and cgroupsCommonly uses OVN databases, ovn-controller, and OVS
Hardware offload depends on eBPF and device capabilitiesEstablished OVS hardware-offload path

The packet paths can be summarized as follows.

Cilium

1
2
3
4
5
6
7
8
9
Pod
 ↓
eBPF
 β”œβ”€β”€ Policy
 β”œβ”€β”€ Service
 β”œβ”€β”€ Routing
 └── Observability
 ↓
NIC

OVN-Kubernetes

1
2
3
4
5
6
7
8
9
Pod
 ↓
OVS
 β”œβ”€β”€ Logical Switch Flow
 β”œβ”€β”€ Logical Router Flow
 β”œβ”€β”€ ACL
 └── Load Balancer Flow
 ↓
NIC

Cilium places more functionality into the Linux kernel’s programmable datapath. OVN places more emphasis on compiling a network-wide logical topology into distributed switch flows.


β‘€ Service, Policy, and Observability

Both platforms provide much more than basic Pod connectivity.

Service Processing

Cilium can replace kube-proxy using eBPF.

1
2
3
4
5
Service
   ↓
eBPF Load Balancer
   ↓
Backend Pod

OVN-Kubernetes represents Kubernetes Services through OVN load-balancer objects and applies them to the relevant logical switches and routers.

1
2
3
4
5
6
7
Service
   ↓
OVN Load Balancer
   ↓
OVS Flow
   ↓
Backend Pod

Both can implement distributed Service processing, but the internal representation differs.

Network Policy

Cilium supports standard Kubernetes NetworkPolicy and Cilium-specific policy resources. Its extended policies can express identity-aware and L7 rules. OVN-Kubernetes supports standard NetworkPolicy as well as cluster-level controls such as AdminNetworkPolicy and OVN-specific egress features.

1
2
3
4
5
Cilium Policy
    ↓
Identity and Endpoint Rules
    ↓
eBPF Policy Maps
1
2
3
4
5
OVN Policy
    ↓
Logical ACL
    ↓
OVS Flow

Observability

Cilium includes Hubble, which derives flow visibility from the eBPF datapath.

1
2
3
4
5
eBPF Flow Events
       ↓
Hubble
       ↓
CLI / Metrics / Service Map

OVN-Kubernetes can expose flow information through OVN sampling and can be combined with platform observability systems such as OpenShift Network Observability. Cilium’s observability is more tightly integrated into its primary networking stack. OVN observability is generally assembled from OVN flow information and surrounding platform components.


β‘₯ Does eBPF Make Cilium Always Faster?

No. Cilium can avoid or replace several traditional Linux networking layers and can accelerate some Service paths through XDP. This can reduce packet-processing overhead in compatible configurations. However, performance depends on more than the name of the CNI. Important variables include:

  • native routing versus overlay tunneling,
  • packet size,
  • MTU,
  • encryption,
  • Service type,
  • conntrack behavior,
  • kernel version,
  • NIC and driver,
  • CPU topology,
  • hardware-offload support.

For example, the following configurations cannot be compared fairly as though only the CNI had changed.

1
2
3
4
Cilium
Native Routing
No Encryption
XDP Acceleration
1
2
3
4
OVN
Geneve Overlay
IPsec
Software OVS

The first configuration naturally has fewer processing steps. Likewise, an OVS flow offloaded to hardware can outperform a host-based eBPF path for the supported traffic, despite OVS having a more layered control plane.

1
2
3
Software eBPF on Host CPU
             vs
OVS Flow in NIC/DPU Hardware

Therefore, the correct statement is:

Cilium can provide a highly efficient software datapath, while OVN can provide a highly efficient hardware-offloaded datapath. The result depends on the actual configuration and workload.


⑦ Why OVN Appears Frequently with DPUs

A DPU is not itself a CNI. It is an infrastructure processor that can execute networking, security, and storage functions outside the host CPU.

1
2
3
4
5
6
7
8
# Without DPU
Pod
 ↓
Host CPU
 ↓
Network Datapath
 ↓
NIC
1
2
3
4
5
6
7
8
# With DPU Offload
Pod
 ↓
DPU / NIC Hardware
 ↓
Network Datapath
 ↓
Physical Network

OVS has an established hardware-offload model. OVS flows can be translated into rules supported by compatible NICs and DPUs. With NVIDIA BlueField, a simplified path is:

1
2
3
4
5
6
7
8
9
OVN-Kubernetes
        ↓
OVS
        ↓
OVS-DOCA or Hardware Offload
        ↓
BlueField DPU
        ↓
Hardware Flow Processing

This integration allows portions of switching, routing, policy, and encapsulation processing to move away from the host CPU. NVIDIA’s DOCA Platform Framework reference deployments explicitly provide an accelerated OVN-Kubernetes service using BlueField DPUs. That makes OVN-Kubernetes a natural option when the primary goal is:

Move the Kubernetes network dataplane from the host CPU to the DPU.

Cilium also has access to Linux hardware-offload interfaces for certain BPF programs. However, Cilium uses multiple eBPF hooks, maps, helpers, socket-level operations, and policy mechanisms. A device must support the required semantics before the entire datapath can be moved into hardware.

1
2
3
BPF Hardware Offload Exists
              β‰ 
Entire Cilium Datapath Is Offloaded

As a result, OVN and OVS currently have a clearer end-to-end product path for BlueField network offloading.


β‘§ Is the Choice Cilium or DPU + OVN?

Not necessarily. The following simplified choice is useful for initial understanding.

1
2
3
4
5
6
7
General Kubernetes
        ↓
      Cilium

DPU-Accelerated Virtual Network
        ↓
    OVN + OVS

However, real clusters can use hybrid designs. An AI cluster may use Cilium for the default Kubernetes interface and a separate high-performance interface for GPU communication.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
Pod
 β”œβ”€β”€ eth0
 β”‚     ↓
 β”‚   Cilium
 β”‚     ↓
 β”‚   Kubernetes Service / API Traffic
 β”‚
 └── net1
       ↓
     Multus
       ↓
     SR-IOV / RDMA
       ↓
     ConnectX or BlueField

In this design:

  • Cilium manages ordinary Kubernetes connectivity.
  • Hubble provides observability for the default network.
  • Multus attaches an additional interface.
  • SR-IOV reduces software-path overhead.
  • RoCE or InfiniBand transports GPU communication.
  • A DPU can accelerate infrastructure functions independently.

This is often more relevant to AI and HPC clusters than forcing all traffic through a single CNI datapath. The practical choices are therefore closer to these three models.

ArchitecturePrimary Use
Cilium on ordinary NICsGeneral cloud-native Kubernetes
OVN-Kubernetes with OVS/DPU offloadOpenShift, virtualization, multitenancy, infrastructure offload
Cilium plus Multus, SR-IOV, or DPUKubernetes control traffic plus high-performance secondary networks

⑨ The Future Is Not Simply eBPF vs OVS

Cilium and OVN began from different architectural assumptions.

1
2
Cilium
Linux Kernel as Programmable Dataplane
1
2
OVN
Logical Network Compiled into Distributed Switches

Their future directions are also different.

Cilium’s Direction

Cilium is likely to continue expanding the Linux kernel as a unified cloud-native dataplane.

1
2
3
4
5
6
7
8
Networking
Security
Service Load Balancing
Observability
Service Mesh
Runtime Enforcement
        ↓
      eBPF

Its strongest vision is reducing the number of disconnected network and security components that must be installed around Kubernetes.

OVN’s Direction

OVN-Kubernetes is likely to continue expanding its logical-network capabilities.

1
2
3
4
5
6
7
8
9
Logical Switching
Logical Routing
Multitenancy
User-Defined Networks
Multiple Interfaces
BGP Integration
Hardware Offload
        ↓
    OVN and OVS

Its strongest vision is providing a programmable virtual network that can span containers, virtual machines, physical networks, and hardware accelerators.

DPU’s Direction

DPU adoption changes the question. The traditional question was:

Which software datapath should run on the host?

The emerging question is:

Which networking functions should remain in the host kernel, and which should move into the NIC or DPU?

1
2
3
4
5
6
7
8
9
10
11
Host Kernel
 β”œβ”€β”€ Workload-aware policy
 β”œβ”€β”€ Socket-level visibility
 └── Application context

DPU / SmartNIC
 β”œβ”€β”€ Switching
 β”œβ”€β”€ Routing
 β”œβ”€β”€ Encapsulation
 β”œβ”€β”€ Encryption
 └── Infrastructure isolation

This suggests that the long-term architecture may not have one universal winner. Cilium can remain strong for workload-aware control inside the host, while OVN and OVS remain strong for virtual network abstraction and hardware offloading.


Key Takeaway

Cilium and OVN-Kubernetes implement overlapping Kubernetes networking functions, but they optimize different architectural layers.

CiliumOVN-Kubernetes
Programs the Linux kernel with eBPFPrograms a distributed virtual network through OVN and OVS
Strong cloud-native integrationStrong logical-network abstraction
Integrated networking, policy, Service, and Hubble observabilityIntegrated switching, routing, ACL, egress, and multitenancy
Natural choice for general KubernetesNatural choice for OpenShift and virtualized infrastructure
Primarily optimized as a host software datapathHas a mature OVS hardware-offload path
Can coexist with SR-IOV and DPUsCan move major dataplane functions onto DPUs

The choice should not be reduced to new eBPF versus old OVS.

A more useful distinction is:

Cilium treats the Linux kernel as the cloud-native network platform. OVN treats the entire cluster as a programmable logical network.

For general Kubernetes, Cilium is often the simpler and more integrated choice. For OpenShift, complex virtual networking, multitenancy, or BlueField-based infrastructure offloading, OVN-Kubernetes has a clearer architectural fit. For AI and HPC environments, the likely result is a hybrid design in which the default CNI manages Kubernetes traffic while SR-IOV, RDMA, and DPUs handle specialized high-performance paths.


TL;DR

Q. Are Cilium and OVN-Kubernetes both CNI solutions?

Yes. Both can provide Pod connectivity, Service processing, and network policy, but their internal architectures are different.

Q. What is Cilium’s main architectural idea?

Cilium programs networking, security, and observability logic into Linux kernel eBPF hooks.

Q. What is OVN-Kubernetes’ main architectural idea?

OVN-Kubernetes models logical switches, routers, ACLs, and load balancers, then compiles them into OVS flows on each node.

Q. Is Cilium always faster because it uses eBPF?

No. Performance depends on routing mode, overlay, encryption, Service configuration, kernel, NIC, packet size, and hardware offloading.

Q. Why is OVN frequently associated with DPUs?

OVS has a mature flow-offload model, and NVIDIA provides reference deployments for accelerating OVN-Kubernetes with OVS and BlueField DPUs.

Q. Can Cilium be used with a DPU?

Yes. Cilium can remain the default CNI while a DPU handles SR-IOV, RDMA, encryption, switching, or other infrastructure functions.

Q. Is the main choice Cilium or OVN with a DPU?

That is a useful simplification, but not a strict rule. Hybrid architectures are common and often more appropriate.

Q. Which should be selected for a general Kubernetes cluster?

Cilium is often the more natural choice when kube-proxy replacement, eBPF policy, and integrated Hubble observability are priorities.

Q. Which should be selected for OpenShift or BlueField offloading?

OVN-Kubernetes is generally the more established option for OpenShift, logical virtual networking, and OVS-based DPU offloading.

Q. What is the likely future?

Cilium will continue expanding kernel-based cloud-native networking, while OVN will continue strengthening logical networks and hardware offloading. AI and HPC clusters will increasingly combine a default CNI with separate accelerated networks rather than relying on one datapath for every traffic type.


References

This post is licensed under CC BY 4.0 by the author.