Docker Compose DNS Fails in Prod: 4 Fixes That Work

Disclosure: As an Amazon Associate, I earn from qualifying purchases. Some links in this post are affiliate links — they cost you nothing extra.
⚡ Key Takeaways
  • Docker Compose DNS works locally but fails in production due to differences between Docker Desktop (VM-based) and Docker Engine on bare Linux servers.
  • Four fixes: explicitly define networks in docker-compose.yml, override host DNS settings with dns_search: [], use external networks with aliases for multi-file setups, and configure /etc/docker/daemon.json to prevent systemd-resolved conflicts.
  • Check /etc/resolv.conf inside containers with 'docker exec' to diagnose DNS issues—look for 127.0.0.11 as the first nameserver and no search domains.
  • Intermittent DNS failures often come from host search domains causing wildcard DNS matches to wrong IPs—clear them with dns_search: [] in your Compose file.

Docker Compose worked perfectly on my laptop. Then I deployed to staging and watched services time out trying to reach each other.

The error was always the same: getaddrinfo ENOTFOUND db or curl: (6) Could not resolve host: api. Containers were running. Logs showed nothing useful. The network existed. But DNS resolution between services just… stopped working.

This isn’t a beginner mistake. Docker Compose networking behaves differently across Docker Engine versions, host network configurations, and especially between Docker Desktop (which you probably use locally) and Docker Engine on a bare Linux server. The DNS resolver that “just works” on your Mac can silently fail in production for reasons that aren’t obvious from the docs.

Here’s what actually breaks, and the four fixes that solved it across three different production incidents.

Close-up of a person holding a Git sticker, emphasizing software development.
Photo by RealToughCandy.com on Pexels

How Docker Compose DNS Resolution Works (When It Works)

Docker Compose creates a bridge network for your services. Each service gets a container name as its hostname. Docker’s embedded DNS server (running at 127.0.0.11 inside each container) resolves service names to container IPs.

When you run docker-compose up, Compose:

  1. Creates a network named {project}_default (or whatever you specify)
  2. Assigns each service container to that network
  3. Registers each service name with Docker’s internal DNS
  4. Configures each container’s /etc/resolv.conf to use nameserver 127.0.0.11

So when your web service tries to connect to postgres://db:5432, the embedded DNS resolver looks up db in its registry, finds the container IP (something like 172.18.0.3), and returns it.

This works flawlessly in development. Then you deploy.

Enjoying this article? Get more like it delivered to your inbox. Subscribe to the newsletter

Fix #1: Explicitly Define the Network (Docker Engine 20.10+ Regression)

The first production incident happened after a Docker Engine upgrade from 19.03 to 20.10.12. Services that had been running for months suddenly couldn’t resolve each other.

The issue: Docker Compose v2 changed how it handles network creation when you don’t explicitly define networks. In some configurations (especially with docker compose instead of docker-compose), the auto-generated network doesn’t properly register service names with the DNS resolver.

The fix is to explicitly define the network in your docker-compose.yml:

version: '3.8'

services:
  web:
    image: myapp:latest
    networks:
      - app_network
    depends_on:
      - db

  db:
    image: postgres:14
    networks:
      - app_network

networks:
  app_network:
    driver: bridge

This forces Docker to create a predictable network with proper DNS registration. Without the explicit networks: block, Docker Compose v2 sometimes creates networks with broken DNS configs—especially when you’ve got multiple Compose files or override files in play.

After this change, DNS resolution started working again. No other changes needed.

Fix #2: Check Your /etc/resolv.conf on the Host

The second incident was weirder. DNS worked for about 80% of requests, then randomly failed. docker exec web ping db would succeed three times, then time out, then succeed again.

The culprit was the host’s /etc/resolv.conf. Our Ubuntu 22.04 server was using systemd-resolved with a search domain configured:

# /etc/resolv.conf on the host
nameserver 127.0.0.53
search internal.company.com

Docker copies this config into containers. The problem: when DNS resolution fails for a bare hostname like db, the resolver appends the search domain and tries db.internal.company.com. If your internal DNS server has a wildcard A record (ours did), it returns a bogus IP. The connection times out. Docker retries. Sometimes it works, sometimes it doesn’t.

The fix: override DNS settings in docker-compose.yml:

services:
  web:
    image: myapp:latest
    dns:
      - 127.0.0.11  # Docker's embedded DNS
    dns_search: []   # Clear search domains
    networks:
      - app_network

Setting dns_search: [] prevents Docker from copying the host’s search domains into the container. Now db resolves to the container IP every time, no random failures.

This is particularly nasty because it’s nondeterministic. You might deploy, see everything work for 10 minutes, then start seeing intermittent connection failures. Logs show connection refused or timeout, not DNS errors, because the DNS “succeeded”—it just returned the wrong IP.

Top view of a book, coffee, and cookies on a bed with white roses.
Photo by Şehâdet Yoldaç on Pexels

Fix #3: Use FQDN Service Names with Network Aliases

Third incident: a microservices stack with 12 services. Some services could reach each other, others couldn’t. No pattern. Logs were useless.

Turns out we had services across multiple Compose files (one for core services, one for workers, one for monitoring). Each file created its own network. Services in different networks can’t resolve each other’s names, even if you manually attach containers to multiple networks later.

The fix: use a shared external network and fully-qualified service names.

# core-services.yml
version: '3.8'

services:
  api:
    image: api:latest
    networks:
      shared_net:
        aliases:
          - api.prod

  db:
    image: postgres:14
    networks:
      shared_net:
        aliases:
          - db.prod

networks:
  shared_net:
    external: true
# workers.yml
version: '3.8'

services:
  worker:
    image: worker:latest
    environment:
      - DB_HOST=db.prod
      - API_HOST=api.prod
    networks:
      shared_net:

networks:
  shared_net:
    external: true

Create the shared network once:

docker network create shared_net

Now all services can resolve each other using the aliases (db.prod, api.prod). This also makes it explicit which services are meant to communicate across boundaries.

The alternative—trying to attach services to multiple auto-generated Compose networks—doesn’t reliably register DNS entries. I’m not entirely sure why Docker’s DNS registration is flaky when you manually attach containers to networks after creation, but using a pre-created external network with explicit aliases has been 100% reliable.

Fix #4: Disable User-Defined Bridge for systemd Conflicts

Fourth incident: DNS resolution worked fine until we enabled a VPN on the production host. After that, containers couldn’t resolve anything—not service names, not external domains, nothing.

The issue: systemd-networkd and systemd-resolved were fighting with Docker’s bridge network. Docker creates iptables rules for its bridge networks, but systemd-resolved was intercepting DNS queries and routing them incorrectly. The symptom: nslookup google.com inside a container would hang for 30 seconds, then fail.

The fix: configure Docker daemon to use iptables legacy mode and explicitly set DNS servers:

// /etc/docker/daemon.json
{
  "iptables": true,
  "dns": ["8.8.8.8", "8.8.4.4"],
  "default-address-pools": [
    {"base": "172.20.0.0/16", "size": 24}
  ]
}

Restart Docker:

sudo systemctl restart docker

Then update your docker-compose.yml to use external DNS for internet traffic but keep Docker’s embedded DNS for service resolution:

services:
  web:
    image: myapp:latest
    dns:
      - 127.0.0.11  # Docker's embedded DNS (for service names)
      - 8.8.8.8     # Fallback for external domains
    networks:
      - app_network

This tells Docker: try embedded DNS first (which resolves service names), then fall back to Google DNS for everything else. The default-address-pools setting in daemon.json prevents IP range conflicts with the VPN.

After this change, both internal service resolution and external DNS worked reliably.

The One Diagnostic Command That Actually Helps

When DNS breaks, don’t waste time with docker logs. Instead, exec into a failing container and check its DNS config:

docker exec -it <container> cat /etc/resolv.conf

You should see:

nameserver 127.0.0.11
options ndots:0

If you see:
– nameserver 127.0.0.53: Docker copied the host’s systemd-resolved config (bad)
– search <domain>: Docker copied search domains from the host (can cause failures)
– Multiple nameservers but 127.0.0.11 isn’t first: Docker’s embedded DNS won’t be tried first

Then test DNS resolution manually:

docker exec -it <container> nslookup db
docker exec -it <container> ping db

If nslookup works but ping fails, the issue is network routing, not DNS. If both fail, it’s DNS misconfiguration.

One more thing: check if the Docker network actually exists and has the services attached:

docker network inspect <network_name>

Look for the Containers section. Every service should be listed. If a service is missing, it’s not attached to the network—check your docker-compose.yml for typos in network names.

When to Just Use IP Addresses Instead

DNS resolution adds a lookup step. For latency-critical services (like a metrics collector that queries a database every 100ms), you might want to skip DNS entirely.

You can assign static IPs in docker-compose.yml:

services:
  db:
    image: postgres:14
    networks:
      app_network:
        ipv4_address: 172.20.0.10

  web:
    image: myapp:latest
    environment:
      - DB_HOST=172.20.0.10
    networks:
      - app_network

networks:
  app_network:
    driver: bridge
    ipam:
      config:
        - subnet: 172.20.0.0/24

Now web connects directly to 172.20.0.10 with no DNS lookup. The downside: if you recreate db, you have to ensure it gets the same IP. And it’s way less readable in logs.

I’d only do this if you’ve profiled DNS as a bottleneck. Otherwise, stick with service names.

Why Docker Desktop Hides These Issues

If you’re developing on a Mac or Windows machine with Docker Desktop, you’ve probably never seen these problems. That’s because Docker Desktop runs Docker Engine inside a lightweight VM with a pre-configured network stack. The VM’s DNS resolver is isolated from your host OS, so systemd-resolved conflicts don’t exist. Network creation is handled by Docker Desktop’s daemon, which uses consistent defaults.

But on a bare Linux server (which is what you’re running in production), Docker Engine interacts directly with the host’s network stack, iptables, and DNS resolver. That’s where things break.

The gap between “works on my machine” and “works in production” is mostly this: Docker Desktop’s VM abstracts away host networking quirks, while Docker Engine on Linux exposes them. You won’t catch these issues until you deploy—or until you test in a Linux VM locally. A KVM-based local testing setup can save you from surprises like these.

FAQ

Q: Why does docker-compose work but docker compose (without hyphen) break DNS?

The hyphenated docker-compose is the standalone Python tool (v1), while docker compose is the Go-based plugin (v2) bundled with Docker Engine 20.10+. They handle network creation slightly differently. v2 sometimes skips DNS registration when networks are auto-generated without explicit definitions. Always define networks explicitly if you’re using v2.

Q: Can I use links: instead of networks to fix DNS?

No. links: is deprecated and only works with the legacy default bridge network. It doesn’t work with user-defined networks, which is what Compose creates. Use networks: and service names.

Q: Should I use 127.0.0.11 or 8.8.8.8 as the primary DNS server in containers?

127.0.0.11 should always be first. That’s Docker’s embedded DNS resolver, which handles service name lookups. Add 8.8.8.8 (or your preferred public DNS) as a fallback for external domains. If you put 8.8.8.8 first, service names won’t resolve because Google’s DNS doesn’t know about your Docker containers.

Use Explicit Networks and Clear DNS Configs

Every production DNS failure I’ve dealt with came down to implicit behavior. Docker Compose auto-generates networks with names you didn’t choose. Docker copies DNS settings from the host without asking. Services across multiple Compose files can’t see each other because networks don’t align.

The fix is to make everything explicit:

  1. Define networks in docker-compose.yml with the networks: block
  2. Set dns: [127.0.0.11] and dns_search: [] to avoid host DNS pollution
  3. Use external networks and aliases for cross-Compose-file communication
  4. Configure /etc/docker/daemon.json if systemd-resolved is interfering

Do this before you deploy, and you’ll skip the 2am debugging session where you’re staring at connection timeouts with no useful logs.

One thing I haven’t tested yet: how IPv6-enabled Docker networks handle DNS. Most production setups are still IPv4-only, but if you’re running dual-stack, the DNS resolver behavior might be different. Let me know if you’ve hit issues there.

Did you find this helpful?

Your support keeps this blog running and ad-free content coming.

☕ Buy me a coffee
TODAY 1,974 | TOTAL 130,181