Migrating Cassandra and Kafka off ingress-nginx: what ingress2gateway misses

September 29, 2026

Migrating Cassandra and Kafka off ingress-nginx: what ingress2gateway misses

September 29, 2026
Migrating Cassandra and Kafka off ingress-nginx: what ingress2gateway misses

TL;DR

  • ingress-nginx was archived read only on 24 March 2026. No releases, no CVE fixes, no exceptions.
  • ingress2gateway does a decent job on HTTP Ingress objects. It cannot see your tcp-services ConfigMap, because those ports were never Ingress objects to begin with.
  • Gateway API has caught up on layer 4. TLSRoute is GA in the Standard channel since v1.5.0, and TCPRoute and UDPRoute since v1.6.0, but some controllers still gate them behind experimental CRDs, optional CRDs or a flag.
  • The hard part is not the YAML. It is Kafka advertised.listeners and long lived CQL sessions, which quietly make a DNS weighted cutover unsafe.

ingress-nginx is archived, and the migration guides stop at HTTP

The Kubernetes project archived ingress-nginx on 24 March 2026. The repository has been read only since, with no releases and no security patches, for a controller that used to sit in front of a large share of the clusters out there. We have been working through this with customers all year.

If you run stateless HTTP workloads, the migration guides already published by AWS, Datadog, Pulumi and others will get you most of the way. Take an inventory, run ingress2gateway, pick a controller, run both side by side, move DNS. That path is well documented and I am not going to repeat it here.

What none of those guides cover is the thing we actually get called about. Our customers run Cassandra, Kafka, ClickHouse and Schema Registry behind that ingress, and those are the workloads where the tooling stops helping you. This post is about the last twenty per cent.

ingress2gateway cannot see your tcp-services ConfigMap

ingress2gateway reached 1.0 in March 2026, and the 1.1 and 1.2 releases since have kept adding to it. Its ingress-nginx provider converts a longer list of annotations than people expect, including ssl-passthrough, rewrite-target, proxy-body-size, the three proxy-*-timeout annotations, the CORS family, whitelist-source-range and the canary annotations. Run it, read the output, and you will have a reasonable starting point for your HTTP estate.

The problem is that your data platform ports were never expressed as Ingress objects. Cassandra CQL on 9042, Kafka broker listeners, the ClickHouse native protocol on 9000: all of those were exposed through the ingress-nginx --tcp-services-configmap flag, which is controller configuration rather than Kubernetes API objects. ingress2gateway reads Ingress resources, so a ConfigMap full of 9042: cassandra/cassandra-0:9042 entries is invisible to it. Nothing warns you. The tool completes successfully and your database ports simply are not in the output.

The first thing I do on any of these migrations is dump that ConfigMap, its UDP sibling and the real annotation inventory, and treat them as separate, hand written work.

kubectl -n ingress-nginx get cm tcp-services -o yaml
kubectl -n ingress-nginx get cm udp-services -o yaml
kubectl get ingress -A -o json | jq -r '.items[].metadata.annotations | keys[]' | sort -u

Gateway API now covers layer 4, but controller support still varies

The good news is that layer 4 is no longer the weak spot it was a year ago. TLSRoute is GA in the Standard channel as gateway.networking.k8s.io/v1 since Gateway API v1.5.0, and TCPRoute and UDPRoute graduated to Standard in v1.6.0, released on 30 June 2026. That same release moved genuinely experimental resources to the gateway.networking.x-k8s.io group with an X prefix, which makes it much easier to tell what you are depending on.

Controller support is where you still have to check rather than assume. At the time of writing, Envoy Gateway, Cilium, NGINX Gateway Fabric and Traefik are all listed as conformant against v1.6.1, but they differ on the layer 4 resources:

Layer 4 route support by Gateway API implementation, at the time of writing
ImplementationTLSRouteTCPRouteUDPRouteWhat to watch
Cilium 1.20SupportedSupportedSupportedTCPRoute and UDPRoute CRDs are optional, and Cilium disables support for them if they are not installed
Traefik 3.7Supported, Standard channelExperimental channel onlyNot listedTCPRoute needs experimentalChannel enabled and the Experimental channel CRDs installed. Enabling the flag without those CRDs stops the whole Gateway provider from starting
NGINX Gateway Fabric 2.7Supported, v1Supported, v1Supported, v1From 2.7 all three are served from the Standard channel with no flag. Older releases need experimental features enabled
Envoy Gateway 1.9Supported, v1Supported, v1Supported, v1Install the v1.6 CRDs first, because without them TCP and UDP routes are silently skipped. Docs examples still use v1alpha2, which is deprecated but still served

TLSRoute when there is SNI to read, TCPRoute when there is not

The choice between TLSRoute and TCPRoute comes down to one question: is there a TLS handshake for the Gateway to read a server name from? Kafka clients send SNI, so many brokers can share a single port on a single Gateway and be told apart by hostname.

Cassandra is more interesting. The CQL protocol carries no hostname of its own, so an unencrypted cluster gives the Gateway nothing to match on and you are back to one TCP listener port per node. Turn on client to node encryption and it changes: the TLS handshake carries the server name, and this is exactly how k8ssandra exposes CQL through Traefik on port 9142, with several clusters behind one port and token aware routing intact. It needs a driver that sets a per node SNI name in the handshake, which the modern Java driver does. If your clients are older, or the cluster is not encrypted, use TCPRoute and a listener per node.

Unencrypted Cassandra needs one listener and one TCPRoute per node

That per node listener needs its own external port, because the Gateway has nothing else to tell the nodes apart by. Here are two nodes, each exposed on its own port and forwarded to plain CQL on 9042 behind the Gateway:

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: data-gateway
  namespace: cassandra
spec:
  gatewayClassName: cilium
  listeners:
    - name: cassandra-0
      protocol: TCP
      port: 19042
      allowedRoutes:
        kinds:
          - kind: TCPRoute
    - name: cassandra-1
      protocol: TCP
      port: 19043
      allowedRoutes:
        kinds:
          - kind: TCPRoute
---
apiVersion: gateway.networking.k8s.io/v1
kind: TCPRoute
metadata:
  name: cassandra-0
  namespace: cassandra
spec:
  parentRefs:
    - name: data-gateway
      sectionName: cassandra-0
  rules:
    - backendRefs:
        - name: cassandra-0
          port: 9042
---
apiVersion: gateway.networking.k8s.io/v1
kind: TCPRoute
metadata:
  name: cassandra-1
  namespace: cassandra
spec:
  parentRefs:
    - name: data-gateway
      sectionName: cassandra-1
  rules:
    - backendRefs:
        - name: cassandra-1
          port: 9042

Yes, that is one listener and one route per node, and yes, it is tedious for a thirty node cluster. Generate it. The alternative is exposing a single service in front of Cassandra, which defeats the token aware routing that makes your driver fast in the first place. Remember that each node must also advertise the address and port clients reach it on through the Gateway, otherwise the driver will discover peers it cannot connect to.

See it working first

If you would rather see that working before you trust it, we put the whole thing in a repository: digitalis-io/ingress2gateway-demo builds a kind cluster with Envoy Gateway in front of a single Cassandra node and runs cqlsh through the TCPRoute to prove the path. It takes about five minutes, most of which is Cassandra starting up.

ClickHouse, Schema Registry and UDP follow the same rules

ClickHouse is two problems. Its HTTP interface on 8123 is usually already an Ingress object, so ingress2gateway converts it with the rest of your HTTP estate. The native protocol on 9000 is the Cassandra problem again: no hostname in the protocol, so without TLS it is a TCPRoute and a listener per replica or per shard you want clients to reach directly. With TLS on the secure native port, 9440, SNI gives you the TLSRoute option instead.

Schema Registry is plain HTTP and is the easy one. It comes across with ingress2gateway, but check the timeouts and body size limits on the new path, because large schemas and slow clients were often tuned with annotations.

Anything in the udp-services ConfigMap, typically DNS, syslog or metrics collectors, maps to UDPRoute, which is in the Standard channel since v1.6.0. Check the table above before you plan around it: Traefik does not list UDPRoute at all, and Cilium only enables it if the CRD is installed.

Kafka will tell your clients the wrong address

This is the one that catches people out, and it has nothing to do with Gateway API.

A Kafka client connects to a bootstrap address, gets back metadata containing advertised.listeners for every broker, and then connects directly to those addresses. If the advertised address still points at the ingress-nginx load balancer while you are weighting DNS towards the new Gateway, your producers happily resolve the bootstrap to the new IP and are then sent straight back to the old one. Half a cutover is worse than no cutover.

Add a cluster-ip listener with advertised hosts alongside the old one

Strimzi has already deprecated its type: ingress listener precisely because ingress-nginx is archived. The path that works today is a cluster-ip listener with explicitly advertised hosts, with the Gateway resources managed by you. Add it alongside the existing listener rather than changing the existing one in place:

listeners:
  - name: external
    port: 9094
    type: ingress
    tls: true
    # existing configuration unchanged, removed at the end of the migration
  - name: gateway
    port: 9095
    type: cluster-ip
    tls: true
    configuration:
      brokers:
        - broker: 0
          advertisedHost: broker-0.kafka.example.com
          advertisedPort: 443
        - broker: 1
          advertisedHost: broker-1.kafka.example.com
          advertisedPort: 443

Each advertised host then gets a TLSRoute matching that hostname, with the Gateway listener in passthrough mode so the broker terminates TLS itself and client certificates still work. Because the broker terminates TLS, the new hostnames must be in the broker certificate SANs. If you supply your own listener certificate, add them before you start.

Strimzi proposal 136 describes a native type: tlsroute listener with advertisedHostTemplate and advertisedPortTemplate fields. It leaves TCPRoute out and expects you to bring your own Gateway. Until it ships in a release you can use, do it by hand.

Move clients to new names, then remove the old listener

The safe sequence treats each listener change as a planned change of its own, because any change to the listener configuration makes Strimzi roll the brokers:

  1. Add the new listener. This is the first rolling restart of the brokers.
  2. Create the Gateway, the TLSRoutes and the DNS names for the new hosts, pointing straight at the Gateway.
  3. Test against the new bootstrap address from a client outside the cluster and confirm the metadata returns the new hostnames.
  4. Move clients to the new bootstrap address, service by service. The old listener keeps serving everyone who has not moved.
  5. Once nothing uses the old listener, remove it. This is the second rolling restart of the brokers.

You are moving clients to names that already work end to end, not swapping the address underneath them.

Annotations fall into three buckets, and NGINX snippets do not convert

Once you have the annotation inventory from earlier, most entries fall into three buckets.

Converted cleanly: timeouts, body size, CORS, redirects, source range allow lists, canary weights and ssl-passthrough. ingress2gateway handles these and the result is readable.

Needs a policy CRD: rate limiting and session affinity cookies have no core Gateway API equivalent, so you are writing a controller specific policy resource. Envoy Gateway has BackendTrafficPolicy and SecurityPolicy, Traefik has its middlewares. These work, but they are no longer portable between controllers, which is worth knowing before you claim the migration made you vendor neutral.

You lose it: configuration-snippet and auth-snippet are raw NGINX configuration and there is nothing to convert them to. In practice this is the section where a migration stalls, because someone pasted twelve lines of Lua into an annotation three years ago and nobody remembers why. Find those early.

DNS weighting is not a cutover plan for long lived connections

For an HTTP service, weighting DNS between two load balancers is a reasonable cutover. For data platforms it is not, for two reasons.

The first is that clients cache resolved addresses. A Cassandra driver opens its connection pool at startup, keeps those sockets for the lifetime of the process and never re-resolves. A JVM with a badly configured networkaddress.cache.ttl will hold a resolved IP effectively forever. Weighting DNS to ten per cent does nothing to a client that resolved the name at eight o'clock this morning.

The second is draining. When you eventually remove the old ingress, every one of those pooled connections is cut at once, and a large Cassandra or Kafka client fleet reconnecting simultaneously is its own kind of outage. Plan a rolling restart of the client applications as the final step of the migration, rather than assuming connections will migrate on their own. It feels crude. It is also the only thing that reliably works.

While you are there, check the equivalent of proxy-read-timeout on the new path. CQL sessions and Kafka consumer connections sit idle for long periods, and a Gateway with a thirty second idle timeout will tear them down in a way that ingress-nginx, with your bespoke annotation, did not.

Start with the inventory, and treat data platforms as a separate project

Clone the demo repository if you want a working TCPRoute in front of a database to poke at. Then pull the tcp-services and udp-services ConfigMaps and the annotation inventory from your own cluster before you touch anything else, because those commands tell you how much of this post applies to you. Run ingress2gateway for the HTTP estate, pick a controller whose layer 4 support matches what you actually need, and treat Cassandra, Kafka and ClickHouse as a separate project with its own cutover plan and a client restart at the end.

Frequently asked questions

Does ingress2gateway migrate the ingress-nginx tcp-services ConfigMap?

No. ingress2gateway reads Ingress resources, and ports exposed through the --tcp-services-configmap flag were never Ingress objects. The tool completes without a warning and the Cassandra, Kafka and ClickHouse ports are simply missing from its output. Dump the tcp-services and udp-services ConfigMaps and rebuild them by hand as TCPRoute, TLSRoute or UDPRoute resources.

Should I use TCPRoute or TLSRoute for Cassandra?

It depends on whether there is a TLS handshake for the Gateway to read a server name from. Unencrypted CQL carries no hostname, so you need a TCPRoute and a dedicated listener port per node. With client to node encryption and a driver that sets a per node SNI name, such as the modern Java driver, you can use TLSRoute and put several nodes or clusters behind one port with token aware routing intact.

Which Gateway API controllers support TCPRoute and UDPRoute?

At the time of writing, Cilium, NGINX Gateway Fabric 2.7 and Envoy Gateway 1.9 support TLSRoute, TCPRoute and UDPRoute. Cilium needs the optional CRDs installed, and Envoy Gateway needs the v1.6 CRDs or routes are silently skipped. Traefik 3.7 supports TLSRoute, offers TCPRoute only through its experimental channel and does not list UDPRoute.

How do you move Strimzi Kafka off ingress-nginx?

Add a second cluster-ip listener with explicitly advertised hosts alongside the existing type: ingress listener. Give each advertised host a TLSRoute with the Gateway listener in passthrough mode, and add the new hostnames to the broker certificate SANs. Test from outside the cluster, move clients to the new bootstrap address service by service, then remove the old listener. Each listener change rolls the brokers.

Why is a DNS weighted cutover unsafe for Cassandra and Kafka?

Kafka clients connect to whatever brokers advertise in advertised.listeners, so a new bootstrap IP can still send them back to the old load balancer. Cassandra drivers open their connection pool at startup and never re-resolve DNS, so weighting has no effect on running clients. Plan a rolling restart of client applications as the last step instead.

Which ingress-nginx annotations cannot be converted to Gateway API?

configuration-snippet and auth-snippet contain raw NGINX configuration and have no equivalent. Rate limiting and session affinity cookies have no core Gateway API equivalent and need a controller specific policy resource, such as Envoy Gateway's BackendTrafficPolicy or Traefik middlewares, which are not portable between controllers.

Do not find out about advertised.listeners in the cutover window

This is the sort of work our Kubernetes and data platform team does week in, week out. We run these migrations on live clusters with the customer's own clients in front of them.

If your Cassandra or Kafka ports still sit behind ingress-nginx, get in touch at digitalis.io/contact.

References and related reading

Subscribe to newsletter

Subscribe to receive the latest blog posts to your inbox every week.

By subscribing you agree to with our Privacy Policy.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Ready to Transform 

Your Business?