Skip to content

Network Issues

genesluna edited this page Mar 15, 2026 · 5 revisions

🌐 Network Issues

This page covers issues related to NetworkPolicies, ingress, health checks, and connectivity problems in PR environments.


⚠️ 502 Bad Gateway - Port Mismatch

Symptoms

  • App pod is running and healthy
  • Health endpoint works internally (kubectl exec ... curl localhost:8080)
  • External access returns 502 Bad Gateway
  • Traefik logs show no errors

Root Cause

The allow-ingress-controller NetworkPolicy uses the configured port from k8s-ee.yaml. If your app listens on a different port than configured, traffic is blocked.

Common Port Configurations

Stack Default Port k8s-ee.yaml Configuration
Node.js/NestJS 3000 Default, no config needed
.NET/ASP.NET Core 8080 app.port: 8080
Go 8080 app.port: 8080
Java/Spring Boot 8080 app.port: 8080
Python/FastAPI 8000 app.port: 8000

Diagnosis

# Check network policy port
kubectl get networkpolicy allow-ingress-controller -n {namespace} -o yaml | grep -A5 ports

# Compare with your app's actual listening port
kubectl exec -n {namespace} {pod-name} -- netstat -tlnp

Resolution

Ensure your k8s-ee.yaml has the correct port configured:

app:
  port: 8080  # Match your application's listening port

Then redeploy (push to PR or re-run workflow) to update the NetworkPolicy.


⛔ Traffic Blocked by NetworkPolicy

Symptoms

  • App running but not accessible from outside
  • Timeout when accessing preview URL
  • Inter-pod communication failing
  • Database connections refused despite pods running

Diagnosis

# List network policies in namespace
kubectl get networkpolicies -n k8s-ee-pr-{number}

# Describe policies for details
kubectl describe networkpolicy -n k8s-ee-pr-{number}

# Test from inside cluster
kubectl run debug --rm -it --image=busybox -n k8s-ee-pr-{number} -- \
  wget -qO- --timeout=5 http://k8s-ee-pr-{number}-demo-app:80/api/health

Common Causes

Cause Symptoms Solution
Missing ingress rule External access blocked Check Traefik namespace selector
Wrong port in policy Some ports blocked Verify app.port in k8s-ee.yaml matches app
Egress blocked DNS or DB connections fail Check egress policy for DNS
Missing pod selector Wrong pods affected Verify label selectors

Understanding NetworkPolicies

PR environments have these network policies applied:

Policy Purpose
default-deny-all Denies all traffic by default
allow-ingress-from-traefik Allows traffic from Traefik ingress
allow-dns-egress Allows DNS resolution
allow-internal Allows pod-to-pod within namespace
allow-observability Allows Prometheus scraping

Resolution

Check if Traefik can reach your app:

# Verify Traefik namespace selector in policy
kubectl get networkpolicy allow-ingress-from-traefik -n k8s-ee-pr-{number} -o yaml | grep -A10 namespaceSelector

# Expected: should match kube-system namespace where Traefik runs

Check if app can reach database:

# Test connectivity to PostgreSQL
kubectl run debug --rm -it --image=postgres:16 -n k8s-ee-pr-{number} -- \
  pg_isready -h k8s-ee-pr-{number}-postgresql-rw -p 5432 -U app

# Test connectivity to MongoDB
kubectl run debug --rm -it --image=mongo:7-jammy -n k8s-ee-pr-{number} -- \
  mongosh --host app-mongodb-0.app-mongodb-svc --eval "db.adminCommand('ping')"

Check DNS resolution:

kubectl run debug --rm -it --image=busybox -n k8s-ee-pr-{number} -- \
  nslookup k8s-ee-pr-{number}-postgresql-rw.k8s-ee-pr-{number}.svc.cluster.local

🚪 Ingress Not Working

Symptoms

  • 404 or 503 when accessing preview URL
  • TLS certificate errors
  • DNS resolution works but HTTP fails

Diagnosis

# Check ingress resource exists
kubectl get ingress -n k8s-ee-pr-{number}

# Describe ingress for details
kubectl describe ingress -n k8s-ee-pr-{number}

# Check Traefik logs
kubectl logs -n kube-system -l app.kubernetes.io/name=traefik --tail=50

# Verify DNS resolution
nslookup k8s-ee-pr-{number}.k8s-ee.genesluna.dev

Common Causes

Cause Error Solution
Ingress not created 404 Not Found Check Helm deployment
Wrong service name 503 Service Unavailable Verify backend service
TLS misconfigured SSL errors Check certificate
Traefik not running Connection refused Check Traefik pods

Expected Ingress Configuration

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: app-ingress
  namespace: k8s-ee-pr-42
  annotations:
    traefik.ingress.kubernetes.io/router.entrypoints: websecure
    traefik.ingress.kubernetes.io/router.tls: "true"
spec:
  rules:
  - host: k8s-ee-pr-42.k8s-ee.genesluna.dev
    http:
      paths:
      - path: /
        pathType: Prefix
        backend:
          service:
            name: app-service
            port:
              number: 80

Resolution

Force Traefik to reload configuration:

kubectl rollout restart deployment -n kube-system traefik

Check service endpoints:

# Verify service has endpoints
kubectl get endpoints -n k8s-ee-pr-{number}

# If empty, pods aren't matching service selector
kubectl describe svc -n k8s-ee-pr-{number} <service-name>

Test backend directly:

# Port-forward to bypass ingress
kubectl port-forward -n k8s-ee-pr-{number} svc/app-service 8080:80

# Test in another terminal
curl http://localhost:8080/api/health

🔒 TLS Certificate Issues

Symptoms

  • SSL_ERROR in browser
  • certificate has expired error
  • unable to verify certificate error

Diagnosis

# Check certificate status (if using cert-manager)
kubectl get certificates -n k8s-ee-pr-{number}

# Check certificate details
kubectl describe certificate -n k8s-ee-pr-{number}

# Test TLS directly
openssl s_client -connect k8s-ee-pr-{number}.k8s-ee.genesluna.dev:443 -servername k8s-ee-pr-{number}.k8s-ee.genesluna.dev

Resolution

If using wildcard certificate: The platform uses a wildcard certificate for *.k8s-ee.genesluna.dev. If TLS fails:

# Check Traefik TLS configuration
kubectl get secret -n kube-system | grep tls

# Verify wildcard secret exists
kubectl get secret wildcard-tls -n kube-system -o yaml

💗 Health Check Failures

💡 Health endpoint requirement: Your app must return HTTP 200 at the configured healthPath. Any response body is accepted — there is no requirement for a specific JSON format like {"status": "ok"}.

Startup Probe Fails

Symptoms:

  • Pod killed before becoming ready
  • Events show Startup probe failed
  • Container restarts with FailedStartupProbe

Diagnosis:

kubectl describe pod -n k8s-ee-pr-{number} <pod-name> | grep -A10 "Startup:"

Common Causes:

Framework Typical Startup Time
Node.js/Bun 1-5 seconds
.NET 5-15 seconds
Java/Spring Boot (ARM64) 60-90 seconds
Java with cold JVM 90-120 seconds

The default startup probe allows 125 seconds (failureThreshold=60, periodSeconds=2).

Resolution:

For apps that need more than 125 seconds, increase the tolerance:

startupProbe:
  failureThreshold: 90  # Gives ~185 seconds
  periodSeconds: 2

Liveness Probe Fails

Symptoms:

  • Pod restarts after running for a while
  • Events show Liveness probe failed
  • Application appears to hang

Common Causes:

Cause Solution
Resource exhaustion Increase CPU/memory limits
Deadlock Check for blocking operations
Slow endpoint Increase timeout
Database connection pool exhausted Increase pool size

Diagnosis:

# Check resource usage
kubectl top pod -n k8s-ee-pr-{number}

# Check probe endpoint manually
kubectl exec -n k8s-ee-pr-{number} <pod-name> -- wget -qO- http://localhost:3000/api/health

Readiness Probe Fails

Symptoms:

  • Pod running but not receiving traffic
  • Service endpoints empty
  • 503 errors from ingress

Diagnosis:

# Check endpoints
kubectl get endpoints -n k8s-ee-pr-{number}

# Check readiness status
kubectl get pods -n k8s-ee-pr-{number} -o wide

# Check pod readiness conditions
kubectl describe pod -n k8s-ee-pr-{number} <pod-name> | grep -A10 "Conditions:"

Resolution:

If readiness probe is checking database connectivity:

# Verify database is accessible
kubectl exec -n k8s-ee-pr-{number} <pod-name> -- wget -qO- http://localhost:3000/api/health/db

📈 Metrics Issues

Metrics Not Appearing in Prometheus

Symptoms:

  • Dashboard panels show "No Data"
  • Prometheus queries return empty results
  • ServiceMonitor exists but metrics missing

Diagnosis:

# Check ServiceMonitor exists
kubectl get servicemonitor -n k8s-ee-pr-{number}

# Check Prometheus target status
kubectl port-forward -n observability svc/prometheus-kube-prometheus-prometheus 9090:9090
# Visit localhost:9090/targets and look for your namespace

# Check if app is exposing metrics
curl https://k8s-ee-pr-{number}.k8s-ee.genesluna.dev/metrics

Common Causes:

Cause Solution
ServiceMonitor missing label Ensure release: prometheus label exists
App not exposing /metrics Check MetricsModule is imported
Port mismatch ServiceMonitor port must match service port name
Network policy blocking Verify observability namespace can reach app
metrics.enabled not set Verify metrics.enabled: true in k8s-ee.yaml

High Cardinality Metrics

Symptoms:

  • Prometheus memory usage increasing
  • Queries becoming slow
  • "cardinality" warnings in Prometheus logs

Common Causes:

Cause Solution
Unique IDs in route labels Middleware normalizes paths
Static asset paths Static assets excluded from metrics
UUID paths UUIDs normalized to :uuid placeholder

How Route Normalization Works:

Original Path Normalized Path
/api/simulator/status/500 /api/simulator/status/:code
/api/simulator/latency/slow /api/simulator/latency/:preset
/user/123 /user/:id
/item/550e8400-e29b-... /item/:uuid
/assets/index-D4IGy2yB.css (excluded from metrics)

🔔 Alert Issues

Alerts Not Triggering

Symptoms:

  • Alert demo running but no alerts fire
  • Dashboard shows no error rate or latency spikes

Diagnosis:

# Check alert demo is running
curl https://k8s-ee-pr-{number}.k8s-ee.genesluna.dev/api/simulator/alert-demo/status

# Check if PrometheusRule exists
kubectl get prometheusrule -n observability custom-alerts

# Port-forward Prometheus to check rules
kubectl port-forward -n observability svc/prometheus-kube-prometheus-prometheus 9090:9090
# Visit localhost:9090/rules

Expected Timeline for Alerts:

Phase Duration What Happens
Start 0s Demo begins sending requests
Metrics scraped 30s Prometheus collects first data points
Rate calculation ~2m Prometheus has enough data for rate calculation
Alert pending ~2m Alert condition becomes true, enters pending state
Alert fires ~7m After 5m in pending state, alert fires
Demo ends 10m 30s Demo stops automatically

💡 Note: Alerts require rate(...[5m]) plus for: 5m pending duration before firing (~10 minutes total).


🔗 Related Pages

Clone this wiki locally