Every machine in an Active Directory domain finds its domain controller by asking DNS. Not from a configured list. It asks for an SRV record, and it goes wherever the answer sends it. Which makes a handful of records under _msdcs the most security-critical thing in the zone, and raises two questions worth answering properly: how do you sign them, and who is allowed to write them. Neither gets asked often, and both get answered badly by default.

This post answers both for a Samba4 domain controller. BIND with dlz_bind9 serving the directory, inline signing, a hidden primary that no client ever reaches, a delegation inside a public zone so the trust chain runs to the root, and client dynamic updates kept well away from the locators.

But signing is only worth anything if you know what it protects, and that means starting one level down — at how a name gets resolved at all. So the first third of this is the walk down from the root. If you already have that cold, skip to Samba’s two DNS backends.

A Resolver Starts Knowing Almost Nothing

It is worth being clear about how little a resolver is born with, because everything else in this post follows from it.

A freshly installed recursive resolver knows two things. It knows the addresses of the root servers — a hints file, thirteen names from a.root-servers.net to m.root-servers.net, served by anycast from a great many more instances than thirteen. And, if it validates, it knows one public key: the root’s.

That is the whole of the built-in configuration. Every other fact it will ever serve — which servers are authoritative for uk, where your domain lives, what address your mail server has — is learned at runtime by asking, and then cached until its TTL runs out.

This is a good design. Nobody has to ship a list of the internet’s name servers, and no central party has to approve a change to your zone. But it has a consequence that the rest of this post is about: a resolver believes what it is told, by whichever server the previous answer pointed it at. Take away signatures and the whole structure is a chain of assertions, each one authenticated only by the fact that it arrived from the address the last answer named. That is a lot of weight to hang on a return address.

The Walk Down the Tree

A lookup is not one question. It is a series of referrals down the name tree, and each step is a separate query to a different server.

Say a client wants dc01.ad.example.co.uk. A resolver with a cold cache does this:

  1. Ask a root server. The root does not know the answer and does not pretend to. It returns a referral: an empty answer section, and in the authority section the NS records for uk, with the addresses of those name servers as glue in the additional section.
  2. Ask a uk server. Another referral, this time to the name servers for example.co.uk.
  3. Ask an example.co.uk server. If ad.example.co.uk is delegated — and in the design later in this post it is — one more referral.
  4. Ask an ad.example.co.uk server. This one is authoritative for the name, so it answers with the A record and sets the AA (authoritative answer) bit.
A single lookup walked down the delegation tree, one referral at a timeRecursive resolverBorn knowing only:· the root hints· the root’s public key1   Root servera–m.root-servers.netReferralthe NS records for uk. — plus their addresses as glue2   uk name serverauthoritative for uk.Referralthe NS records for example.co.uk.3   example.co.ukholds the delegationReferralad.example.co.uk. is delegated — ask its servers4   ad.example.co.ukauthoritative for the nameAnswerthe A record for dc01, with the AA bit setEvery step is cached until its TTL expires, so a warmresolver starts at step 3 or 4.Which is exactly what makes one poisoned cache entryworth so much.
One lookup, four servers. Each step returns not the answer but the name of somewhere better to ask, and the resolver caches every step on the way down.

Three things about that walk matter later.

The referral is the delegation. A zone’s parent does not hold its contents. It holds NS records saying “ask over there”, and — once signing enters the picture — a DS record saying “and here is the fingerprint of the key you should expect when you do”. Delegation is the only structural mechanism DNS has, and it is the thing you will use to keep an internal zone internal.

Nearly all of this is cached. The root and TLD referrals have long TTLs, so a warm resolver skips straight to step three or four. This is why a resolver’s cache is worth attacking: poison one entry and you have redirected everything under it for as long as the TTL you chose.

The full name is not sent to every server. Under QNAME minimisation, a resolver asks the root only about uk, not about dc01.ad.example.co.uk. Worth knowing if you ever go looking for your internal names in a query log upstream. They should not be there.

Stub, Recursive, Authoritative

Three words that get used interchangeably and mean quite different jobs. The distinction is load-bearing for the Active Directory half of this post, so it is worth nailing down.

What it doesWalks the tree?Holds zone data?
Stub resolverThe library or local service your applications call. Asks one configured server and takes the answer.NoNo
ForwarderPasses queries to another resolver, caches the replies.NoNo
Recursive resolverDoes the walk above, caches every step, optionally validates signatures.YesNo
Authoritative serverAnswers for zones it has been given, and only those. Says “I don’t know” about everything else.NoYes

The important line is the last one. An authoritative server has no business doing recursion, and a recursive resolver has no business being authoritative. Combining them is how you get a machine that will both accept arbitrary questions from clients and hold data those clients depend on — which is exactly the machine you do not want your domain controller to be.

On a modern Linux desktop there is another layer worth being aware of. systemd-resolved runs a stub listener on the loopback address and writes a resolv.conf pointing at itself, with options edns0 trust-ad. That trust-ad is the one to notice: it tells the stub to believe the AD (authenticated data) bit in replies from that server. The AD bit is not proof of anything on its own — it is the upstream resolver’s claim that it validated. Trusting it is reasonable when the upstream is yours and the hop to it is trustworthy, and meaningless otherwise.

How Services Are Actually Found

An A record answers one question: what address does this name have. It says nothing about which port the service is on, which of several servers to prefer, or what to do when the first one is down.

SRV records answer all three. From RFC 2782, the shape is:

_service._proto.name.   TTL   IN   SRV   priority weight port target
The parts of an SRV record and what each one decidesthe transportthe domainthe servicewhich kind of server_ldap._tcp.dc._msdcs.ad.example.com.600INSRV0100389dc01.ad.example.com.one owner name — underscores keep it out of host spaceTTLrecord typepriorityweightporttargetLower priority is tried first. Weight shares loadbetween the servers that sit at the same priority.The target must be a host name with A or AAAA records —never a CNAME.A single target of "." means the service isdeliberately not offered here.
An SRV record dissected. The owner name encodes the service and the transport; the record body carries the failover order, the load share within an order, the port, and a target that must be a real host name.

Each field earns its place:

  • The underscore prefixes keep the service labels in a namespace of their own. _tcp can never collide with a host actually called tcp, because a host name may not begin with an underscore.
  • Priority is the failover order — lower is tried first. Servers at a higher priority number are only used when everything below is unreachable.
  • Weight shares load within one priority. A pair at weight 100 and weight 300 gets roughly a quarter and three quarters of the clients. It is a proportional draw, not round-robin.
  • Port frees the service from a well-known number, which is how a client finds an LDAP server on 3268 without anyone hard-coding it.
  • Target must be a name with A or AAAA records. RFC 2782 is explicit that it must not be a CNAME, and that a single target of . means “this service is deliberately not offered here” — a useful thing to publish on purpose.

This is what makes a domain locatable rather than configured. A domain-joined machine does not have a list of your domain controllers. It has a domain name, and it asks.

The Names Active Directory Publishes

An AD domain is, from a client’s point of view, mostly a set of SRV records. The tree under _msdcs is the interesting part, because it is how clients distinguish “an LDAP server” from “a domain controller for this domain” from “a global catalogue for this forest”:

NameWhat asks for it
_ldap._tcp.<domain>Anything wanting LDAP in the domain
_ldap._tcp.dc._msdcs.<domain>Domain controller location — the big one
_ldap._tcp.pdc._msdcs.<domain>The PDC emulator specifically
_ldap._tcp.gc._msdcs.<forest>Global catalogue
_kerberos._tcp.<domain>, _kerberos._udp.<domain>KDC location, before any ticket exists
_kpasswd._tcp.<domain>, _kpasswd._udp.<domain>Password changes
_ldap._tcp.<site>._sites.dc._msdcs.<domain>Site-aware location — find a DC that is near

Look at what that list is for. Kerberos authentication cannot begin until the client has found a KDC, and it finds the KDC by DNS. Site-aware location means the answer also decides which datacentre a client authenticates against.

The Same Idea, Modernised

SRV has a successor worth knowing about. SVCB and HTTPS records (RFC 9460) generalise the pattern: service parameters in DNS, including ALPN, port, and address hints, in one record. That is how a browser learns to go straight to HTTP/3 without a redirect. The HTTPS record is already widely deployed. The mechanism is the same in spirit as SRV: the client is told, by DNS, how to reach the service. Which means the argument in the next section applies to it just as much.

Everything After the Lookup Trusts the Lookup

Here is the pivot, and it is the reason a DNS post has a security section rather than the other way round.

A domain-joined machine boots and asks where a domain controller is. It gets a name and a port. It connects there, and then it starts doing all the things we normally think of as the security layer: Kerberos, LDAP signing, channel binding, certificate validation.

So what happens if the answer was a lie?

Being fair about this matters, because the answer is not “instant catastrophe”. Kerberos was designed with mutual authentication exactly so that a client is not at the mercy of its name lookup: a host that cannot produce a service ticket for the name the client asked for cannot complete the exchange. Get an SRV answer pointing at a machine with no key in the domain and, for a client configured strictly, it fails.

The realistic damage is subtler, and it is all in the gaps around that:

  • Downgrade. A client that falls back to NTLM when Kerberos does not work has just been handed to whoever answered.
  • Relay and coercion. The attacker does not need to be the DC. Being the name a client connects to is enough to start relaying that authentication somewhere it is useful.
  • Denial of service that looks like a fault. Point _ldap._tcp.dc._msdcs at something that does not answer and the domain is intermittently broken in a way nobody diagnoses as DNS for a good while.
  • Everything with no mutual authentication at all. Time sync, syslog, monitoring, backup agents, that internal HTTP thing with verify=False. Plenty of estates have more of these than they would like to admit.

The general principle is the one worth taking away: DNS is a discovery mechanism, not an authorisation mechanism. It is entirely reasonable to find a service by name. It is not reasonable to grant anything on the strength of a name, and the number of systems that quietly do — PTR-based allowlists, host-name ACLs, “it’s on the internal network so it must be ours” — is the actual attack surface.

What DNSSEC Fixes, and What It Does Not

DNSSEC exists to answer one question: did this answer really come from the zone that owns the name, unmodified?

It works (RFC 4033 and the two that follow it) by signing, and by chaining the signatures to something you already trust:

  • Every RRset in a signed zone has an RRSIG — a signature over that set.
  • The zone’s public keys are published as DNSKEY.
  • The parent zone publishes a DS record: a hash of the child’s key.
  • That DS is itself signed by the parent, whose key is hashed in its parent’s DS, all the way up to the root — and the root’s key is the one thing your resolver was born knowing.
A chain of trust from the root down to an internal zone. — the rootthe one key your resolver already hasWhere every chain starts.Shipped with the resolver, rolled rarely, trusted implicitly.signed DS for com.com.or uk., or whatever you sit underA parent vouches for its child's key.The DS is a hash of the child's key, signed by the parent.signed DS for example.com.example.com.public, signed, DS lodged with the parentThe zone you already own and already sign.It holds the delegation for the internal zone, and the DSthat authenticates it. That is all it holds of the inside.signed DS for ad.example.com.above the line: published to the internetbelow it: served only inside the estatead.example.com.the AD-derived view — internal servers onlySigned inside, trusted from outside.Your resolvers validate it root → com → example.com → here,with no local trust anchor and nothing to distribute.The data never leaves. Only the trust comes in, downthe ordinary public chain.
The chain a validating resolver builds. Each parent vouches for its child’s key with a signed DS record, so one built-in trust anchor at the root authenticates every zone below it — including a delegated internal zone whose contents never leave the building.

Denial is signed too, which is easy to overlook and matters more than it sounds. Without it, “no such name” is an unauthenticated answer an attacker can forge to make something disappear. NSEC and NSEC3 give an authenticated “nothing exists between these two names”.

That is a real fix for a real problem. Cache poisoning is not theoretical: the Kaminsky attack made off-path spoofing cheap enough to force emergency patching across the entire internet, and the mitigations that followed — source port randomisation, 0x20 case mixing — are all attempts to make guessing harder, not to make answers verifiable. Side-channel work since has repeatedly chipped away at them. Signatures are the only thing in that list that changes the game rather than raising the price.

Now the honest limits, because DNSSEC is oversold in both directions.

It is not confidentiality. Signing is public-key authentication; every query and answer is still in clear text on the wire. Privacy on the client hop is DoT, DoH or DoQ, and those are a different mechanism solving a different problem. Encrypting the hop to a resolver that does not validate buys you a private conversation with something that can still be lied to.

It does not validate meaning, only origin. DNSSEC proves the zone’s owner published this. If the record is wrong, or malicious, or was written by someone the zone unwisely allowed to write it, the signature is applied to that just as faithfully. Signing a zone that untrusted machines can write does not make the contents trustworthy — it authenticates the lie. Hold onto that sentence; the last third of this post is essentially its consequence.

It only helps if somebody validates. If your resolver does not validate, signatures on the zones you query are decoration. And if validation happens at a resolver across the network from the client, the client is trusting the AD bit and the path — see trust-ad above.

And it needs operating. An expired signature is not a degraded answer, it is SERVFAIL — the name goes dark. A DS in the parent that no longer matches the child’s key does the same. This is the honest cost, and as such the automation in the next sections matters more than the initial signing.

Why a Made-Up Internal TLD Cannot Be Signed

This is where naming decisions made years ago come back with a bill, and it is worth spelling out because it is the reason this post’s design uses a real, owned domain for internal names.

Look again at how the chain is built: a zone’s key is vouched for by a DS record in its parent. So a validating resolver can only authenticate an internal zone if that zone has a parent willing and able to publish a DS for it.

ad.corp.local has no such parent. Neither does .lan, .home, or anything else invented on the day the domain was provisioned:

  • .local is reserved for mDNS. Using it for unicast DNS is not merely unsigned, it is a documented collision with how every operating system in the building is entitled to behave. It gets its own section below, because it is in a different league to the other three.
  • .lan, .corp, .home are unregistered strings. There is no parent to hold a DS, and the applied-for versions were withheld from delegation because of the name-collision mess in private estates.
  • home.arpa is properly reserved for home networks, which sounds ideal — but its delegation is deliberately insecure. There is no signed path down to it, by design.
  • .internal has been set aside by ICANN for exactly this use, and it solves the collision problem. It does not solve this one: no delegation means no DS, so no chain.

Your options with an unchained internal zone are to leave it unsigned, or to configure a local trust anchor on every validating resolver in the estate and become the root of your own private island — a key you now have to distribute, monitor and roll by hand, on every resolver, forever, with SERVFAIL across the entire domain as the failure mode.

There is a much easier answer, and it is free: make the internal zone a delegation inside a public zone you already own and already sign. ad.example.com, delegated from example.com. The DS goes in the public parent. Every validating resolver in the estate then authenticates internal names with the trust anchor it already has, and you distribute nothing.

The contents stay inside. Only the trust comes from outside. That distinction is the design in the rest of this post.

.local Is Not a Style Choice

One of those four deserves its own section, because it is the one people defend, and because the defence is always the same: it works, we have used it for years, what is the problem.

The problem is that .local is not an unregistered string somebody might one day sell. It is a namespace that already has an owner and a defined behaviour, and the behaviour is not “ask the DNS server”. RFC 6762 §3 is not gentle about it:

Any DNS query for a name ending with “.local.” MUST be sent to the mDNS IPv4 link-local multicast address 224.0.0.251 (or its IPv6 equivalent FF02::FB).

Read what that MUST actually governs. It is not a claim about who owns the string. It is an instruction about where the query goes — and the answer is a multicast group on the local link, not your domain controller. The same section says names under .local are “meaningful only on the link where they originate”, the DNS equivalent of a 169.254 address.

So a client that follows the specification correctly will never ask your DNS server for a .local name. It shouts on the wire and takes whatever answers.

That produces a set of failures with a very particular flavour:

  • Resolution depends on the client, not on your DNS. macOS resolves .local through Bonjour and always has; systemd-resolved routes the local domain to mDNS by default; Windows has done mDNS since Windows 10. Three stacks, three sets of rules, and your zone file is not consulted by any of them.
  • It stops at the first router. mDNS is link-local by design. A name that resolves at a desk on the same VLAN does not resolve from another floor, another site, or the VPN — which is the “works in the office, breaks at home” ticket that gets closed as “network issue” four times before somebody reads the spec.
  • The behaviour changes underneath you. Whether a given machine tries multicast first, unicast first, or both in parallel depends on the resolver stack and its version. Estates that ran on .local “fine for years” are usually estates where a distro update has not yet changed the order.
  • The evidence is missing from where you look for it. The DC’s query log shows nothing, because nothing arrived. People spend days on the server, and the query never left the client.
  • And it cannot be signed, which is the point of this section. No parent, no DS, no chain, no way to publish one.

Now put that under an Active Directory domain. The realm is derived from the domain name. The service principals are derived from the realm. The _msdcs locators — the records this entire post is about — sit beneath a suffix that a conforming client is required to resolve by shouting at the local segment. You are building Kerberos on top of a namespace that half your estate resolves with a protocol designed for finding printers.

And the exit is expensive, which is what makes the original decision worth being unsentimental about. There is no in-place rename on Samba. The documented path is samba-tool domain backup rename followed by a restore: you take a renamed copy of the database, re-seed a new DC from it, and re-add every other DC from scratch, with no overlap permitted between old-name and new-name DCs. It is not Windows’ rendom, which at least walks the DCs one at a time. It is a rebuild with a nicer name.

So the honest summary. .local for a unicast domain is not a preference, a convention, or a harmless bit of legacy. It is a documented conflict with a protocol that ships enabled on every operating system in the building, and the bill arrives years later as intermittent, unreproducible resolution failures — and, when you finally want to sign the zone, as a domain rebuild.

The Same Mistake, Wearing a Different Hat

While we are here: port 5353.

mDNS listens on UDP 5353, and it is not a spare port that happens to be free. Standing a normal unicast DNS service on it, or pointing resolvers at :5353 because 53 was busy or needed root, does the same damage from the other direction — every mDNS-aware machine on that segment is now talking to your name service, and your name service is now fielding multicast service discovery it was never meant to answer. The name and the port are two halves of one namespace, and both of them already belong to somebody else.

.local on 53, or unicast DNS on 5353. Same misunderstanding, same class of intermittent failure, same weeks of somebody else’s time.

And Be Honest About What That Tells You

Microsoft recommended .local in the Small Business Server era, which is why so many estates still carry it, and stopped recommending it a long time ago. Inheriting a .local domain is bad luck. Most people reading this who have one did not choose it, and the section above is a migration plan rather than an accusation.

Deploying one is a different matter entirely, and I am going to be blunt about it, because softening this has not helped anybody.

Anybody who deploys .local for Active Directory or for standard unicast DNS, or who stands a DNS service on port 5353, is not an IT professional. They may hold the job title. They are not doing the job. RFC 6762 has been published since 2013, it is readable in twenty minutes, and the sentence that settles the entire question is in section 3. Building a company’s name service on a namespace the specification reserves for link-local multicast is not a defensible engineering decision. It is somebody guessing in the one part of the stack where guessing produces failures nobody can reproduce and everybody blames on the network.

They should be removed from your IT department. Not moved sideways, not given DNS to look after with supervision. Removed from the function. The role exists to know which packet goes where and on whose authority. Somebody who has not read the document governing the namespace they chose for every machine in the building is not performing that role, and keeping them in it means the next decision of this size gets made the same way.

And every system they touched should be audited. This is the part people skip, and it is the part that matters most. A decision like this is never isolated. Somebody who did not check RFC 6762 before naming the domain did not check anything else either — so go and look at what else they built. Expect to find:

  • DNS servers open to the world, forwarders answering anybody who asks, and no ACLs anywhere.
  • Dynamic update left wide open, and zone containers with permissions nobody has reviewed since the domain was created.
  • Certificates and Kerberos built on top of names that were never going to resolve consistently, with the failures papered over by hosts files.
  • Hosts files. Everywhere. Whole estates have been held together by them exactly because .local never worked properly and somebody found a workaround instead of a cause.
  • Firewall rules and service accounts created to make the symptoms go away, still in place, still granting more than anybody now remembers.

That is not vindictiveness. It is what a competence signal is for. When you find one decision that was made without reading the spec, the correct response is to assume the rest were made the same way and go and check — because the same person configured your authentication, your certificates and your access control, and you now have direct evidence of how they approach a problem they do not fully understand.

Inheriting this mess costs you a migration. Employing the person who is still creating it costs you considerably more.

Samba’s Two DNS Backends

A Samba AD domain controller stores its DNS data in the directory itself, and there are two ways to serve it.

SAMBA_INTERNAL is Samba’s own DNS server, built into the AD DC. It handles the AD zones and Kerberos-authenticated dynamic updates, and it hands anything else to a forwarder. Samba describe it as supporting “the basic feature required in an AD” and recommend it “for simple DNS setups”, which is fair and worth taking literally. It gets its own section below, because what it does not do is longer and more interesting than what it does.

BIND9_DLZ runs BIND as the DNS server, with Samba’s dlz_bind9 module loaded into it. DLZ — Dynamically Loadable Zones — is a BIND interface for backing a zone with something other than a zone file. The module answers BIND’s queries out of sam.ldb directly, so there is no copy, no export step, and no synchronisation to go wrong: BIND is reading the directory as it serves.

The Internal DNS Server Does Not Do Recursion — It Cannot

Start with the thing that is almost always described wrongly, including by people who run it.

“The DC is our DNS server, it does recursion for the clients” is not what is happening, because the internal DNS server cannot do recursion at all. Samba’s own feature list says it plainly. The internal DNS does not support:

  • acting as a caching resolver
  • recursive queries (but it can forward to another recursive DNS nameserver)
  • shared-key transaction signature (TSIG)
  • stub zones
  • zone transfers
  • Round Robin load balancing among DCs

with scavenging and conditional forwarders listed as not implemented either.

Read the first two together, because that is the whole story. It cannot resolve, and it cannot cache. What dns forwarder actually buys you is a relay: a client asks the DC for windowsupdate.com, the DC asks a real resolver, the answer comes back through the DC, and then the DC forgets it entirely. The next client asks the same question and the whole thing happens again.

Look back at the table earlier in this post and note what that is: a forwarder without the forwarder’s cache. It has the role’s costs — an extra hop, a dependency, a thing to be down — and none of its benefit.

So a DC with SAMBA_INTERNAL serving the estate is not a DNS server in the sense people mean. It is an uncached proxy for a real resolver, and you have placed it on the machine that holds your directory.

Why That Is a Risk and Not Just an Inefficiency

The inefficiency is easy to see: every external lookup in the building becomes a round trip through the AD DC, permanently, with no cache to take the edge off. On a quiet domain nobody notices. That is exactly why it survives.

The security argument is the one worth making, and it has three parts.

The listener is in the wrong process. The internal DNS is a service task of the AD DC itself, running with the directory’s privileges — not a separate daemon under its own account the way named is. So the thing parsing unauthenticated UDP from anything that can reach port 53 is running inside the process that serves LDAP and Kerberos and owns sam.ldb. BIND has thirty years of hostile attention, a dedicated user, and a habit of being run in a jail, exactly because a DNS listener is a rough neighbourhood. Samba’s internal server is a convenience feature that happens to live in the crown jewels.

Serving clients means being reachable by clients. To be the estate’s DNS server, it must accept queries from every workstation, every printer, every contractor’s laptop on the guest VLAN somebody bridged by accident. That is a large, permanently open, unauthenticated attack surface on the most valuable host you own, and you are running it to save the cost of a resolver that a Raspberry Pi could host.

And a forwarder open to the world is somebody else’s weapon. A DC forwarding for anything that asks is an open forwarder. Once it is reachable off your network it becomes a reflection and amplification participant, which means the traffic and the abuse reports both arrive at your domain controller. There is no rate limiting to reach for, because the internal server has none.

None of that requires a Samba vulnerability to be a bad idea. It is a bad idea on shape alone: it puts an unauthenticated, internet-facing-by-accident, parser-heavy service in the same process as your directory, to do a job it is documented as not being able to do properly.

It Will Fall Over Sooner, and Take More With It

Worth being clear about what this argument is not. It is not a claim that Samba’s DNS code has more bugs than BIND’s. BIND has a long CVE list, mostly because it is the most examined DNS implementation in existence, and counting advisories would be a poor way to choose between them.

The comparison that matters is structural, and it comes down to three questions with three uncomfortable answers.

Who can send it a malformed packet? With SAMBA_INTERNAL serving the estate: every workstation, every phone on the wireless, everything that can route to port 53 on that box. With the design in this post: the signer. One host, one TSIG key, allow-query limited to it. That is not a small difference in degree. It is the difference between an exposed service and one that is effectively unreachable, and it dwarfs any difference in code quality between the two implementations.

What falls over when it does? This is the one that decides the severity. named is a separate daemon under its own account; when it dies, DNS stops and the domain controller carries on authenticating. Samba’s internal DNS is a service task inside the AD DC, so anything that wedges, exhausts or crashes it is happening inside the process that serves LDAP and Kerberos. A DNS problem becomes a directory outage. And systemctl restart named costs a second, whereas restarting a DC is a different sort of morning.

What can you do about it while it is happening? BIND has response rate limiting, allow-query, allow-recursion, blackhole, per-view policy, and the option of not answering clients at all. The internal server has dns forwarder and a log file. When something starts hammering it, there is no dial to turn.

Then add the missing cache. Every client query is a fresh outbound round trip, so a query flood costs the DC an upstream lookup per packet rather than a cache hit — and it costs it in the same process that is trying to issue Kerberos tickets. You do not need an exploit for that; you need a busy morning, a misbehaving application, or somebody pointing a scanner at the wrong VLAN. Sustained, it presents as authentication being slow and nobody thinking to look at DNS.

So yes — even with DLZ in the picture, BIND is the safer place for this. The DLZ module does give named access to Samba’s data, and that is a real consideration this post has already made a point of. But a named crash is a DNS outage rather than a directory outage, and in this design that named is not taking queries from the estate in the first place. Bugs are a fact of every codebase. Blast radius and reachability are things you choose.

And It Cannot Build the Design in This Post

There is a simpler, more final reason it is not the backend here.

No zone transfers. The pipeline in the next section — hidden primary, signer, standalone authoritative servers — starts with an AXFR out of the DC. SAMBA_INTERNAL has nothing to transfer with. No TSIG, either, so even the authentication that transfer would need is absent. And no signing, and no validation of anything it forwards.

So the honest summary of SAMBA_INTERNAL: it is the backend that lets you stand up an AD domain in a lab on a Sunday afternoon without configuring BIND, and it is very good at that. Samba say “simple DNS setups” and they mean it. It is not a resolver, it was never built to be the DNS service for an estate, and the moment you want signing, transfers, views, ACLs, caching or rate limiting, the answer is not to tune it. It does not have those dials. The answer is BIND.

If you are running it today with every client pointed at the DC, the fix is not urgent but it is not optional either: give the clients a real validating resolver, and set dns forwarder on the DC to point at that, so the DC is answering for its own zones and nothing else.

And this is the part I want to be clear about, because “don’t use DLZ” gets repeated as though it were a hardening rule: DLZ is not the exposure. It is the extraction mechanism. What matters is not which module is loaded into BIND — it is who is allowed to talk to that BIND, and what happens to the zone next. A DLZ back-ended BIND that answers only a transfer request from your signer is not an attack surface in any interesting sense. A SAMBA_INTERNAL DC fielding every name lookup from four hundred laptops very much is. Only one of those two turns up on hardening checklists, and it is not the one that matters.

Two things about DLZ are true and worth planning around rather than fearing:

  • The module is version-coupled to BIND. Samba ships a separate .so per BIND version, and named.conf names one specifically. A BIND major upgrade means the matching module must be in place, or named does not start. It is a packaging dependency to test in advance, not a security property.
  • named needs access to Samba’s data, which is why Samba keeps a dedicated directory for the bits BIND needs rather than granting it the whole private directory. The grant is meant to be narrow — worth verifying it still is on your DCs, since it is the one place where DLZ does widen what a compromise of named would reach.

The reason DLZ is the right choice here is what it enables: BIND’s DLZ interface supports enumerating an entire zone, which is what makes a zone transfer out of a DLZ-backed zone possible at all. That transfer is the first hop of the pipeline, and it is BIND doing it, which means the rest of the pipeline is ordinary BIND configuration rather than anything exotic.

There is one constraint that shapes everything downstream, and it is worth stating plainly because it is easy to assume otherwise. A DLZ zone cannot itself be signed. ISC are explicit about it in the BIND ARM: DLZ “is unable to handle DNSSEC-signed data due to its limited API”. You cannot hang a dnssec-policy off the dlz statement and be done.

What you can do — and what ISC suggests for DLZ in the same breath — is run it as a hidden primary, with the signing done by a normal BIND zone that transfers the data in. That is the next section, and the constraint is the reason it has the shape it does.

It is worth knowing that this is a Samba constraint rather than a law of Active Directory. Microsoft’s DNS server has done online signing of dynamic, AD-integrated zones since Windows Server 2012. The zone is signed in place, the private keys replicate to the Key Masters through AD replication itself, and dynamic updates keep working. On Windows, “sign the AD partitions” is a real option and the answer to the churn objection is built in. (On Server 2008 R2 it was not: you could sign an AD-integrated zone, but not one accepting dynamic updates, and every change meant re-signing by hand — which is where the folklore about AD zones being unsignable comes from.)

Samba has no equivalent. Neither backend signs: the internal server has no DNSSEC at all, and DLZ cannot carry signed data. So on Samba the transfer-and-sign pipeline is not one design among several. It is the way.

The Publishing Pipeline

Now the shape of it. Four roles, and the DC is at the back where nothing can reach it.

Publishing an Active Directory zone from a hidden-primary domain controllerDomain controller — hidden primaryBIND with dlz_bind9, authoritative from the directoryrecursion off · transfers to the signer only, with TSIGThe directory host takes no questions.The highest-value machine in the estate is notreachable by the machines that depend on it.AXFR on the SOA refresha poll — DLZ cannot notifySigning instance — a normal secondarytransfers the zone in and serves it signeda second named on the DC, or its own hostThe DLZ zone cannot be signed itself.So signing is a normal secondary holding the copy —and with clients out of the zone, there is no churn.transfer outthe signed zoneAuthoritative serversplain secondaries of the signed zoneno keys, no route to the directoryA compromise gets a copy, not a forgery.With no key on the servers that face the estate, nonew record can be made that will validate.queriesanswered with signaturesValidating resolverswhat client resolv.conf points atinternal zones forwarded, public tree walkedValidation happens next to the client.The chain runs to the root trust anchor, because theinternal zone is a delegation in a public one.no client-facingDNS on the DCPublication runs one way, from the directory outwards.Nothing a client sends ever arrives at the top of it.
The DC is a hidden primary: it serves the AD zone by transfer and answers nothing else. The signing instance is an ordinary secondary zone holding the transferred copy — the DLZ zone itself cannot be signed, so signing always happens one hop along, whether that hop is a second named on the DC or a separate host. The standalone authoritative servers are the only thing clients ever see, and they are secondaries of the signed zone.

1. The domain controller — hidden primary. BIND with dlz_bind9, authoritative for the AD zones out of the directory. Recursion off. No client-facing service. allow-transfer restricted to the signer alone, with TSIG. From the network’s point of view the DC does not serve DNS at all, and the only thing that ever queries it is the next box along.

2. The signing instance — a normal BIND zone that transfers the data in. This is where inline-signing and the dnssec-policy live, on a zone of the same name that holds the transferred copy. BIND keeps the unsigned copy it received and the signed copy it publishes as separate things, and re-signs as new transfers arrive. Because it is a transfer rather than a shared file, this instance can sit on the DC itself — a second named on its own address — or on a separate host, and the configuration is nearly identical either way.

The trade is the one it looks like: on the DC is one fewer machine to run, while a separate host keeps private keys off the box that holds the directory. Both are defensible, and the choice does not change anything else in the pipeline. What matters is that signing happens once, at a defined point, under a key policy — the difference between DNSSEC you operate and DNSSEC that expires at three in the morning.

3. The authoritative servers — the only thing clients see. Plain secondaries of the signed zone. They hold no keys, do no signing, and have no path to the directory. If one is compromised, the attacker has a copy of a zone and no ability to forge a new record that will validate.

4. The resolvers. Validating recursive resolvers for the estate, which conditionally forward the internal zones to those authoritative servers and walk the public tree for everything else. These are what client resolv.conf points at.

Two properties fall out of this arrangement that are worth stating on their own, because they are the whole point:

  • The machine holding the directory is not reachable by the machines using it. A domain controller is the highest-value host in the estate. Giving it a client-facing network service — one that answers unauthenticated UDP from any workstation — is a poor trade for a service that other machines can perform.
  • Signing happens once, at a defined point. The usual objection to signing an AD zone is churn — that the content changes too often for signatures to keep up. That objection is really an objection to the default layout, not to signing. With client registration moved out, as the next section argues it must be, the AD zone changes when a domain controller is promoted or demoted and roughly never otherwise. A zone written only by DCs is a stable zone, and a stable zone is an unremarkable thing to sign. Keeping clients out is not only a security control; it is what makes the zone quiet enough to sign cleanly.

How the Signing Is Actually Wired Up

The shape above is the important part, but “a normal BIND zone that transfers the data in” deserves showing rather than describing, because the first attempt at this usually founders on trying to sign the DLZ directly.

On the DC, the DLZ side stays deliberately dull. It serves the directory, it hands the zone to exactly one peer, and it does nothing else:

key "transfer-to-signer" {
    algorithm hmac-sha256;
    secret "...";
};

options {
    recursion no;
    allow-query { key transfer-to-signer; localhost; };
    allow-transfer { key transfer-to-signer; };
    notify no;
};

dlz "AD DNS Zone" {
    database "dlopen /usr/lib64/samba/bind9/dlz_bind9_18.so";
};

Note the .so carries the BIND major version in its name. That is the version coupling from the previous section made concrete — a BIND upgrade needs the matching Samba module in place before named will start.

On the signing side, the zone is an ordinary secondary with the signing options attached. This is the part that cannot live on the DLZ:

dnssec-policy "ad-internal" {
    keys {
        ksk lifetime P365D algorithm ecdsa256;
        zsk lifetime P90D  algorithm ecdsa256;
    };
};

zone "ad.example.com" {
    type secondary;
    primaries { 192.0.2.10 key transfer-to-signer; };
    file "ad.example.com.axfr";
    inline-signing yes;
    dnssec-policy "ad-internal";
    allow-transfer { key transfer-to-public; };
    also-notify { 192.0.2.20; 192.0.2.21; };
};

That this works on a secondary is the load-bearing detail, and the ARM states it directly:

If yes, BIND 9 maintains a separate signed version of the zone. An unsigned zone is transferred in or loaded from disk and the signed version of the zone is served with, possibly, a different serial number.

So BIND keeps two copies — the unsigned one it received, written to file, and the signed one it serves, written alongside it with a .signed extension. A transfer arrives, the signed version is regenerated, and the downstream servers get a notify, because at this point it is an entirely ordinary zone. inline-signing yes is in fact the default once a dnssec-policy is attached; it is written out above because a config that says what it does is worth the line.

Key rollover comes with the policy rather than with a cron job. ZSK rollovers need no input at all; KSK rollovers need the new DS getting to the parent, which is the CDS/CDNSKEY automation from later in this post. rndc dnssec -status ad.example.com tells you where each key is in its lifetime.

Whether that block runs on the DC or on its own host is a matter of which address primaries points at. On the DC it is a second named instance on a second address, transferring from the first over the loopback or a management interface. That is what “BIND with DLZ can be the signer” means in practice: same software, same box if you like, but the signing happens on the transferred copy rather than on the DLZ zone.

The one operational catch is notify — and it is on the DLZ side. ISC’s manual is blunt about it: DLZ “has no built-in support for DNS notify”, so secondary servers are not automatically informed of changes to the zones in the database. Samba can change a record in the directory and the DLZ instance has no idea it should tell anybody.

So the first hop is a poll, not a push. The signer refreshes on the SOA timer of the zone it is transferring, which means:

  • Propagation from a DC change to a signed, published record is bounded by that refresh interval, not by seconds. Promote a DC and its new _msdcs records appear downstream up to one refresh later.
  • That interval is the dial to turn if the delay matters. It is a trade against how often you want the DLZ queried, which brings up the other thing ISC says about DLZ: it does real-time database lookups with no caching and is “not recommended for use on high-volume servers”.
  • Both of those are arguments for this topology rather than against it. The only client the DLZ instance ever has is the signer, asking once per refresh. The estate’s actual query load lands on the standalone authoritative servers, which are serving a plain signed zone file at full speed.

Everything downstream of the signer is conventional: the authoritative servers are secondaries of the signed zone, they get a proper notify, and they never touch the directory or a key.

Clients Must Not Write the Zone That Holds the Locators

This is the important one, and it is a design rule rather than a setting.

Active Directory registers records dynamically. A machine joins, and it registers itself; a DC starts, and it registers the SRV records that advertise its services. “Secure” dynamic update means those updates are authenticated — the machine proves it is who it says it is with its own credentials, and there is a per-record ACL so a machine can generally only modify a record it created.

Read that carefully, because the guarantee is narrower than it first appears. Secure dynamic update authenticates who is writing. It does not evaluate what the record means. And the set of accounts entitled to write is far wider than people assume: in a default AD-integrated zone the Authenticated Users group holds Create All Child Objects on the zone container in the directory, because ADIDNS stores every record as an AD object under CN=MicrosoftDNS,DC=DomainDnsZones. Not just machine accounts — any authenticated account in the domain, including the one belonging to whoever opened the invoice attachment this morning.

Which produces two failure modes in a zone that holds both client records and service locators:

Names that do not exist yet belong to nobody. Per-record ACLs protect a record that already has an owner. A name that has never been registered has no object to enforce an ACL on, so the first account to create it gets it. That is the mechanism behind the whole family of ADIDNS attacks, the sharpest being a wildcard: create * and every name in the zone that nobody has explicitly claimed — typos, decommissioned hosts, wpad — resolves to the attacker. Existing records are untouched, which is exactly why it goes unnoticed.

The blast radius includes the locators. The records under _msdcs are how every client in the domain finds a domain controller and a KDC. They are the most security-critical records you own, and in a default deployment they sit in the same zone that four hundred laptops write to every time they get a DHCP lease.

Who may write which zone: one combined zone against a split by writerDefault — one zonead.example.comlocators and client records together_ldap._tcp.dc._msdcs.SRV_kerberos._udpSRVdc01Alaptop-042A*A← unclaimeda name nobody has registered has no ACL to enforceevery joined machinemay write all of itOne compromised machine account can rewritethe records that locate a domain controller.Split by who writesad.example.com — DCs onlyno client has an update path into this zone_ldap._tcp.dc._msdcs.SRV_kerberos._udpSRVdc01ANS delegationdyn.ad.example.comor a separate Samba zone with its own ACLlaptop-042Aevery joined machinemay writeno pathtolocators
Who may write what. On the left, one zone, and every joined machine is an authorised writer in the zone that holds the DC locators. On the right, the AD zone is written only by DCs, and client registrations land in a separate sub-zone where the worst a compromised machine can do is lie about itself.

So the rule: the zone that holds the locators is written by domain controllers, and by nothing else. Client dynamic registration goes somewhere else.

Somewhere else can be either of two things, and both are fine:

  • A delegated sub-zone served elsewhere. The AD zone holds an NS delegation for, say, dyn.ad.example.com, and the records land on a separate server that accepts the updates. Samba’s partitions never take a client write at all.
  • A separate zone in Samba with its own update ACL. Still in the directory, but its own zone, so a client write has no path to _msdcs or to a DC’s own records.

The first gives the harder separation; the second is less to run. What fails the rule is neither of those. It is the default, where the two live together. Which is how most domains are still running, because nobody chose it and nobody has been back to look.

And there is a second dividend, the one that ties this back to the pipeline. A zone that only domain controllers write is a zone that hardly ever changes: a DC promotion, a DC demotion, and otherwise silence. All the churn in a default AD zone is client registration. Take that out and the objection to signing the AD partitions goes with it — there is no stream of updates for the signatures to chase, so signing it is routine rather than a fight. The write discipline and the signing are the same decision seen twice.

Alongside either of those, there is a permission to go and look at — and on a Microsoft DC it is the direct fix. Because ADIDNS records are directory objects, the entitlement comes from an AD ACL rather than from anything in the DNS protocol: Authenticated Users holding Create All Child Objects on the zone container. Tightening that is the cleanest mitigation for the unclaimed-name and wildcard problem, and in many estates the permission can be removed outright once you know what actually needs to self-register. Samba’s AD DNS stores its records in the directory the same way, so the same question applies — go and look at what your zone container actually grants.

But Hang On — Should Clients Be Registering At All in 2026?

Everything above assumes client dynamic update is a thing you need and are trying to make safe. Before accepting that, it is worth asking the question nobody asks, because the answer has changed since this behaviour was designed.

Dynamic DNS registration was built for a desktop. A box under a desk, one network cable, one address it kept for years. In that world a machine registering its own name was tidy and basically true.

Now look at what a client is in 2026. It wakes up on home wifi. It comes into the office and joins the corporate wireless. It goes into a docking station and picks up a wired address as well. Somebody starts the VPN and a tunnel adapter appears with a third address. They go to a coffee shop, tether to a phone, and the VPN comes back on a fourth. That is one machine, one name, and half a dozen addresses in a working day — and by default it will try to register a good number of them.

So the zone fills up with claims that were true once.

  • Windows registers every adapter it has, unless somebody has been round and unticked Register this connection’s addresses in DNS per interface. A docked laptop on the VPN is a machine with three live adapters and an opinion about all of them.
  • Multiple A records for one name is not an error state, it is the normal result. A lookup returns them all, clients try them in whatever order they fancy, and connections to that name fail in proportion to how many of the addresses are dead. This is the mechanism behind “remote support can see the machine one minute and not the next”.
  • VPN addresses are the worst of them, because a tunnel address is valid for an hour and the record outlives it. The tunnel drops, the pool address goes to somebody else, and the name now points at a colleague.
  • Docking stations muddy the identity itself. Unless MAC address pass-through is configured, the lease belongs to the dock rather than the laptop, so a hot-desk estate has names, leases and machines drifting apart from each other daily.
  • And ownership makes it permanent. A record can only be updated by the account that created it. When a record was made by DHCP under one credential and the machine later tries to update it under its own, the update fails, quietly, and the stale address stays exactly where it is.

The clean-up story is not the rescue it sounds like either. Samba has had scavenging since 4.9, but it is off by default (dns zone scavenging = yes, with samba-tool dns zoneoptions --aging=1), and Samba themselves say it “should only be enabled on new zones or new installations”, because older versions marked dynamic records as static and static ones as dynamic. On the estates most likely to be full of rubbish — the ones that have been running for years — the tool for clearing it up is the one you are advised not to switch on. It has also had a CVE of its own.

So put the question directly: what actually consumes a laptop’s A record?

In most outfits, very little. Users connect to servers; servers do not connect to laptops. The genuine consumers are remote support tools, RDP to a named workstation, and inventory or monitoring — and nearly all of that tooling maintains its own inventory and works from an agent checking in, because it could never rely on DNS for mobile clients in the first place.

Which suggests inverting the default:

  • Servers and infrastructure get records from provisioning. With NetBox and Ansible already in the picture, the record is created by the same thing that created the machine, it is correct by construction, and it is removed when the machine is.
  • The stable wired estate can take records from DHCP if something actually needs them, with one credential owning them so updates do not fail.
  • Mobile clients register nothing at all. They are consumers of DNS, not publishers of it. If something needs to reach a laptop, it needs an agent, not an A record.

You end up in the same place the security argument put you, from a completely different direction. Fewer writers means a smaller ADIDNS surface, a zone that is not full of expired claims, and — back to the pipeline — a zone quiet enough to sign without thinking about it.

The security case says clients must not write the zone that holds the locators. The operational case asks why they are writing DNS at all. In 2026, for a fleet that changes address five times a day, “they are not” is a perfectly good answer, and considerably less work than making their mess safe.

Split Views, and Where the Trust Comes From

The last piece ties the two halves of the post together.

There are two views of the namespace. A public zone, published to the internet, holding the handful of names the world needs. And an internal view — the AD-derived content, every joined host, every service locator, the site topology — which the world has no business seeing. That content is a map of the estate, and it should be unreachable and untransferable from outside.

But the trust for the internal view comes from the public side, and that is what makes this design better than the usual internal-DNS island:

  • example.com is public and signed, with its DS in the parent and a chain to the root.
  • ad.example.com is delegated from it. The public parent publishes the delegation and a DS for the internal zone’s key.
  • The internal authoritative servers serve the signed ad.example.com. The internal resolvers validate it — root → com → example.com → ad.example.com — using nothing but the root trust anchor they already had.

No local trust anchor. No island. No hand-distributed key. The data never leaves the building, and the validation path is the ordinary public one. When you add a resolver, it validates internal names correctly with no DNSSEC configuration at all.

Being straight about the trade, because there is one. Publishing a delegation and a DS in the public zone means the existence of ad.example.com, and the names of its name servers, are public. The contents are not, and never are — but you have told the world that the zone exists. In exchange, every resolver you own validates internal names against the real root. That is a good trade for most estates, and it should be a deliberate one rather than a surprise.

Two things to get right alongside it:

  • Keep the internal view unenumerable and untransferable. allow-transfer on the internal authoritative servers is for the signer and your own secondaries, nothing else. And bear in mind that NSEC authenticated denial lets anyone who can query the zone walk it end to end; NSEC3 raises that cost, but the real control is that outsiders cannot reach the servers at all.
  • Automate the DS. A DS in the parent that stops matching the child’s key takes the whole internal domain to SERVFAIL. CDS/CDNSKEY exist so the child can signal a key change and the parent can pick it up without a human editing a record during a rollover. If the parent zone is at a registrar or provider that supports it, use it; if not, the rollover procedure needs writing down before the first roll, not during it.

Making Windows and Linux Actually Validate

Everything up to here has been about publishing a zone that can be verified. None of it does anything until something on the client side insists on verifying it. A perfectly signed zone and a client that never checks a signature produce exactly the same experience as an unsigned zone, right up until the day they don’t.

There are only two places validation can happen, and the difference between them is the difference between a security control and a polite suggestion.

Validate at the resolver, and trust the AD bit. The client asks a resolver, the resolver does the cryptography, and it reports the result by setting one bit — AD, authenticated data — in the reply. The client believes the bit. This is the model Windows uses, and it is only as strong as the path between the client and the resolver, because anything that can answer as the resolver can set that bit.

Validate on the client itself. The machine runs its own validating resolver, so the “path to the resolver” is a loopback socket inside the machine and there is nothing left to spoof. This is stronger, and on Linux it is entirely achievable.

The Default Is Nothing

Before configuring anything, it is worth seeing what a mainstream Linux workstation does out of the box. This is a Fedora 44 machine, systemd 259, untouched:

$ resolvectl status | head -3
Global
         Protocols: LLMNR=resolve -mDNS -DNSOverTLS DNSSEC=no/unsupported
  resolv.conf mode: stub

$ grep options /etc/resolv.conf
options edns0 trust-ad

$ dig +dnssec cloudflare.com A | grep flags
;; flags: qr rd ra; QUERY: 1, ANSWER: 3, AUTHORITY: 0, ADDITIONAL: 1

Read those three together, because they tell a small story.

The stub is configured with trust-ad — it has been told to believe the AD bit. systemd-resolved reports DNSSEC=no/unsupported, so it is not validating anything itself. And the reply for a signed zone comes back with flags qr rd ra and no ad — nothing anywhere in that path claimed to have validated.

That is not a misconfiguration. That is the default. A client can be told to trust an assertion that nothing in the chain is making. Nowt is checking it. Worth sitting with for a minute before configuring anything else.

Linux

Three options, in increasing order of how little you have to trust the network.

1. systemd-resolved, validating locally. A drop-in rather than editing the shipped file:

# /etc/systemd/resolved.conf.d/dnssec.conf
[Resolve]
DNSSEC=yes
DNSOverTLS=opportunistic

Then systemctl restart systemd-resolved and confirm with resolvectl status that the line now reads DNSSEC=yes.

The setting to be careful about is the middle one. DNSSEC=allow-downgrade looks like a sensible compromise and is not a security control — resolved.conf(5) says so itself:

Note that this mode makes DNSSEC validation vulnerable to “downgrade” attacks, where an attacker might be able to trigger a downgrade to non-DNSSEC mode by synthesizing a DNS response that suggests DNSSEC was not supported.

An attacker who can forge answers is exactly the attacker DNSSEC exists to stop, so a mode they can switch off by forging an answer buys you nothing against them. It is yes or it is decoration.

2. A real validating resolver on the host. systemd-resolved’s validator is convenient rather than thorough. Where it matters, run Unbound or BIND on the loopback and point the stub at it:

# unbound: validate against the root anchor, refuse to be stripped
server:
    module-config: "validator iterator"
    auto-trust-anchor-file: "/var/lib/unbound/root.key"
    harden-dnssec-stripped: yes
    val-permissive-mode: no

# the internal zone is reached like any other name — no local anchor needed
forward-zone:
    name: "ad.example.com."
    forward-addr: 192.0.2.53

BIND’s equivalent is one line — dnssec-validation auto; — which uses its built-in copy of the root anchor and manages the rollover for you.

3. Estate-wide, at the resolvers you already run. These are stage four of the pipeline earlier in this post. Validation happens there, clients trust the AD bit, and the hop between them is the thing you have to protect — with DoT, or with a network you are willing to make that assumption about.

And note what is not in any of those configurations: a trust anchor for the internal zone. Because ad.example.com is a delegation inside a publicly signed zone, every one of these validates internal names through the ordinary chain from the root. That is the design from the previous section paying for itself. The alternative is pushing a local anchor to every client and resolver in the estate, and re-pushing it at every rollover.

Windows

The important thing first, because it is routinely misunderstood: the Windows DNS Client does not validate DNSSEC. It performs no cryptography, checks no signature, and holds no trust anchor. It is a stub resolver, and it always has been.

What you can do is force it to refuse answers that were not validated by the server on its behalf. That is the Name Resolution Policy Table, and it is per-namespace rather than global:

# Require validated answers for the internal zone
Add-DnsClientNrptRule -Namespace ".ad.example.com" `
    -DnsSecEnable -DnsSecValidationRequired

# What is actually in force on this machine, including from Group Policy
Get-DnsClientNrptPolicy -Effective
Get-DnsClientNrptRule

For the estate, the same thing lives in Group Policy under Computer Configuration → Policies → Windows Settings → Name Resolution Policy: create a rule for the namespace, tick the DNSSEC option, and tick the requirement that the client check the data was validated by the DNS server.

Two things follow from that, and both matter.

Something upstream still has to do the validating. The NRPT rule makes the client demand the AD bit; it does not create one. The resolver those clients point at must be a validating resolver, or every name in that namespace fails.

And this is why the NRPT rule has IPsec options next to it. Microsoft put them there for the reason set out at the top of this section: requiring a bit that any on-path attacker can set is not much of a requirement. If you are relying on the resolver-validates model on an untrusted network, the last hop needs protecting — IPsec between client and resolver, or DoT where the resolver supports it.

It Fails Closed, So Roll It Out in That Order

Forcing validation converts a class of silent compromise into a class of loud outage. That is the correct trade, and it is still an outage: an expired RRSIG, a DS in the parent that no longer matches after a rollover, or a resolver that cannot reach the parent zone all produce SERVFAIL, and SERVFAIL for _ldap._tcp.dc._msdcs means the domain is down rather than degraded.

So do it in this order:

  1. Turn validation on at the resolvers first, and leave the clients alone. Watch for SERVFAIL in the resolver logs for a couple of weeks — this is where you find the zone that has been quietly broken for a year.
  2. Automate the DS before forcing anything, per the previous section. Most self-inflicted DNSSEC outages are a rollover where the parent was never updated.
  3. Then force the clients, one namespace at a time, starting with your own workstation and a test OU rather than the whole estate.

The failure you are engineering against is a client being handed a forged domain controller. The failure you are risking is a client being handed nothing at all. The second one is recoverable and obvious; the first is neither. You will hear about the outage inside a minute. You would never have heard about the other one at all.

Encrypting the Last Hop: DoT and DoH on the Internal Resolvers

The validation section left one thing hanging. In the resolver-validates model — the one Windows gives you — the client is trusting a single bit set by the resolver, and that bit is only worth the path it travelled over. Something has to protect that path.

BIND does support both encrypted transports natively, so this is a configuration job rather than a procurement one:

  • DNS over TLS — a tls block referenced from listen-on, conventionally on port 853.
  • DNS over HTTPS — the same tls block plus an http block, on 443.
  • Outgoing DoT, because forwarders takes a TLS transport per address or for the whole list.
  • Zone transfers over TLS, because a type secondary zone’s primaries statement takes one too — which is directly useful to the pipeline earlier in this post.

First, be clear about what this buys, because DoT and DNSSEC get conflated constantly and they are not alternatives. DNSSEC authenticates the data, all the way back to the zone that published it. DoT protects the conversation with the resolver. One survives a hostile resolver and a hostile network between resolvers; the other stops the machine on your wifi reading and rewriting what your laptop asked. You want both, and neither substitutes for the other. Encrypting the hop to a resolver that does not validate is a private conversation with something that can still be lied to.

Serving It

tls internal-resolver {
    key-file "/etc/pki/dns/resolver.key";
    cert-file "/etc/pki/dns/resolver.pem";
    protocols { TLSv1.3; };
};

http internal-doh {
    endpoints { "/dns-query"; };
};

options {
    dnssec-validation auto;

    listen-on          port  53                          { 192.0.2.53; };
    listen-on          port 853 tls internal-resolver    { 192.0.2.53; };
    listen-on          port 443 tls internal-resolver
                                 http internal-doh       { 192.0.2.53; };
    listen-on-v6       port 853 tls internal-resolver    { 2001:db8::53; };
};

The certificate is the actual work, and it is the part that gets skipped. A client that verifies — which is the entire point — needs a certificate valid for the name it was configured with, issued by something it already trusts. That means your internal CA and your existing certificate automation, not the ephemeral keyword. ephemeral generates a throwaway self-signed certificate; it is there so you can prove the listener works, and it is worthless to any client actually checking.

Forwarding Over It

If those resolvers forward anywhere rather than walking the tree themselves, the upstream hop can be encrypted too — and there is a distinction here worth getting right:

tls upstream {
    ca-file "/etc/pki/tls/certs/ca-bundle.crt";
    remote-hostname "dns.example.net";
};

options {
    forwarders port 853 tls upstream { 192.0.2.1; };
};

Without remote-hostname, you get encryption without authentication: the traffic is unreadable to a passive observer, and an active attacker who can intercept the connection simply presents their own certificate. With remote-hostname and ca-file, BIND verifies who it is talking to. The first is worth something. Only the second is worth calling a control.

The Client Half Is Not Symmetrical

This is where a mixed estate gets awkward, and it is the reason to configure both transports rather than picking one.

Linux does DoT properly. systemd-resolved takes DNSOverTLS=yes for strict mode, and the server can be given the name to verify against:

[Resolve]
DNS=192.0.2.53#resolver.ad.example.com
DNSOverTLS=yes
DNSSEC=yes

As with DNSSEC=, the middle setting is the trap: DNSOverTLS=opportunistic falls back to clear text when TLS is unavailable, which an attacker who can interfere with the connection can arrange.

Windows does DoH, and not DoT. Client DoH support shipped in Windows 11 and Server 2022, configured per server with a template:

$doh = "https://resolver.ad.example.com/dns-query"
netsh dnsclient add encryption server=192.0.2.53 dohtemplate=$doh

DoT, at the time of writing, has only appeared in Insider builds. So on released Windows the encrypted option is DoH or nothing, which is exactly why the NRPT section earlier reached for IPsec instead.

Hence serving both from the same BIND instance. DoT for the Linux fleet and anything else that speaks it, DoH for Windows, one resolver, one certificate.

Which One, Where

For an internal resolver, DoT is the better transport and DoH is the compatibility answer.

DoT sits on its own port. You can see it, allow it, deny it, and alert on anything doing DNS that is not doing it. DoH’s advantage — indistinguishable from ordinary web traffic on 443 — is a genuine benefit on a hostile network and a nuisance on your own, where being able to tell what is DNS is a feature you paid for. On the estate you control, prefer the transport you can observe, and run DoH because Windows leaves you no choice rather than because it is better.

While You Are There: Encrypt the Transfers

The pipeline earlier in this post moves the AD zone by AXFR, and the split-view section made the point that its contents are a map of the estate. TSIG authenticates those transfers; it does not conceal them. Since a secondary’s primaries statement accepts a TLS configuration, the transfer can run over TLS as well — RFC 9103 if you want the standard:

zone "ad.example.com" {
    type secondary;
    primaries { 192.0.2.10 port 853 tls xfr-tls key transfer-to-signer; };
    ...
};

Authenticated by the key, encrypted by the transport. If any hop in that pipeline crosses a site link, a hypervisor you share, or anything you would not happily put a hub on, it is worth the twenty minutes.

What It Does Not Fix

  • It is not validation. Covered above, and worth repeating because vendors sell “secure DNS” meaning encryption alone.
  • It does not hide anything from the resolver. The resolver sees every query in full. Encryption protects the path, not the privacy of the lookup from the operator — which is fine when the operator is you.
  • It does nothing for a client that does not verify the certificate, and opportunistic modes are downgradeable by exactly the attacker you are worried about.
  • And it is not a reason to put a listener on the domain controller. Encrypted DNS on the DC would be solving the wrong problem beautifully. The DC still answers nobody.

How to Check What You Have

Commands to run against your own estate. The outputs are the interesting part, and a couple of them tend to be uncomfortable reading the first time.

# Walk the tree yourself, one delegation at a time
dig +trace dc01.ad.example.com

# Validate, and show the chain being built
delv +rtrace +vtrace ad.example.com SOA

# Is the resolver you are pointed at actually validating?
# A deliberately broken test name must come back SERVFAIL, not an address
dig @<resolver> dnssec-failed.org A

# What does the estate advertise as a domain controller?
dig SRV _ldap._tcp.dc._msdcs.<domain>
dig SRV _kerberos._udp.<domain>

# Is a DC answering for names it has no business answering?
# Ask it for something it is not authoritative for. "recursion requested
# but not available" is the answer you want. An actual address means it is
# serving the estate — recursing if it is BIND, relaying to the forwarder
# if it is SAMBA_INTERNAL. Either way it should not be doing that.
dig @<dc> www.example.org A

# Is the DC configured as the estate's DNS relay?
grep -E 'dns forwarder|server services' /etc/samba/smb.conf

# Will a DC hand its zone to anybody who asks?
dig @<dc> AXFR ad.example.com

# Which backend is this DC running, and does named have the module?
grep -r dlz /etc/named.conf /var/lib/samba/bind-dns/ 2>/dev/null
samba-tool dns query <dc> <domain> @ ALL

# Is the internal zone chained to the public parent?
dig DS ad.example.com @<public-authoritative-for-example.com>

# What does the local stub actually do with the AD bit, and is the
# hop to the resolver encrypted?
resolvectl status                    # DNSSEC= and DNSOverTLS= per link
grep options /etc/resolv.conf        # trust-ad, trusting whom exactly?

# Did anything in the path claim to have validated? Look for "ad" in the flags
dig +dnssec ad.example.com SOA | grep flags

# Validate independently of whatever the local resolver believes
delv ad.example.com SOA              # "fully validated" is the line you want

# Is the resolver actually listening for DoT, and does its certificate
# match the name clients are configured with?
kdig +tls @192.0.2.53 ad.example.com SOA
openssl s_client -connect 192.0.2.53:853 \
    -servername resolver.ad.example.com </dev/null 2>/dev/null \
    | openssl x509 -noout -subject -dates

And on a Windows client, to see whether it is demanding anything at all:

Get-DnsClientNrptPolicy -Effective    # the rules actually in force
Resolve-DnsName ad.example.com -DnssecOk
Get-DnsClientDohServerAddress         # is the hop to the resolver encrypted?

The two that most often produce a surprise are the recursion check and the AXFR attempt. If a DC answers either of them for an arbitrary client, the pipeline in this post is not in place, whatever the diagram on the wiki says.

The third is resolvectl status on a machine nobody has touched. DNSSEC=no/unsupported alongside trust-ad in resolv.conf is the normal state of a Linux desktop, and it means the signing work described above is currently being checked by nobody.

The Short Version

A resolver is born knowing the root servers and one key, and learns everything else by being told. Names are found by walking down delegations, and services are found by asking for an SRV record — so by the time a client starts doing Kerberos with a domain controller, the identity of that domain controller came from a DNS answer. Discovery by DNS is correct and fine. Authorisation by DNS is not, and a surprising amount of infrastructure quietly does it anyway.

DNSSEC is what makes those answers verifiable: signatures on every set, a DS in each parent, a chain to a single trust anchor at the root, and authenticated denial so a name cannot be made to disappear. It buys origin authentication and integrity — not privacy, and not correctness. It signs whatever the zone says, which is why it cannot rescue a zone that untrusted machines are allowed to write. And it needs a parent: an invented internal TLD has nowhere to put a DS, so .local, .lan, .internal and home.arpa all leave you either unsigned or running a private island of hand-distributed keys.

For an Active Directory domain, that produces a design rather than a list of settings. Use a delegation inside a public zone you own, so the trust comes down the ordinary chain from the root while the data never leaves. Serve the AD partitions with BIND and dlz_bind9 — DLZ is not the risk, it is how you get the zone out of the directory — and let the DC be a hidden primary that transfers out and answers nobody else. A DLZ zone cannot itself be signed, so the signing is a normal secondary zone holding the transferred copy, with inline-signing and a dnssec-policy on it: a second named on the DC if you want fewer machines, a separate host if you want private keys off the directory. Either way it is signed once, at a defined point, under a key policy — and the first hop is a poll rather than a push, because DLZ cannot send a notify. Publish from standalone authoritative servers that hold no keys and have no route to the directory.

And keep clients out of the zone that matters. Secure dynamic update authenticates the writer, not the meaning, and in a default AD-integrated zone the writers are Authenticated Users — every account, not just every machine — so a zone holding both laptop records and _msdcs locators is one phished user away from a client being told, with a perfectly valid signature, that the domain controller is somewhere else. Client registrations belong in a sub-zone, delegated out or separate in Samba, where the worst a compromised account can do is lie about itself.

Though the better question is whether clients should be registering at all. A 2026 laptop has an address on home wifi, another on the office wireless, another through the dock and another on the VPN, and it will cheerfully publish most of them. What consumes a laptop’s A record is almost nothing — the tools that need to reach a workstation keep their own inventory, because DNS was never reliable for mobile clients anyway. Records for servers should come from provisioning, and the fleet should register nowt.

Then make something check the signatures, because none of the above is worth anything until a client refuses an answer. On Linux that means DNSSEC=yes in systemd-resolved, or a real validating resolver on the loopback — never allow-downgrade, which an attacker capable of forging answers can simply switch off by forging an answer. On Windows it means accepting that the DNS client never validates anything itself, and using an NRPT rule to make it require an answer the resolver validated, with the last hop protected because that requirement is a single bit. Turn it on at the resolvers first and watch for SERVFAIL, automate the DS, then force the clients. It fails closed, which is the right way round and still an outage.

None of this is exotic. It is delegation, transfer and signing — the three things DNS has always done — arranged so that the machine holding your directory is not the machine taking questions from the car park, and so that when something does lie to a client, the client notices.