fix: dev: Use correct port and target for NOTIFY(CDS)
If there is a DSYNC RRset with multiple records, and unsupported scheme/type records follow supported ones, the port and target of the last record were being used to queue the notify. This does not necessarily match the port and target of the supported record. This has been fixed.
Closes #6080
Merge branch '6080-dsync-mismatch-queue' into 'main'
Only save the DSYNC target and port when a record matches CDS NOTIFY,
then after scanning the complete RRset require count == 1 before using
the stored values.
Current system tests cover a mixed RRset in but the unsupported and
supported records use the same target and port, so the test does not
catch that the wrong port and target are being used for the supported
type.
Change the test such that the supported DSYNC record is followed
by the unsupported ones, and use different port and target for
supported and unsupported DSYNC records.
Mark Andrews [Tue, 14 Jul 2026 23:08:44 +0000 (09:08 +1000)]
fix: test: test-syncplugin treats firstlbl as a prefix, not an exact label
The label length and the string length where not being checked so
a label that started with a string that matched the skip label would
incorrectly match.
Closes #6212
Merge branch '6212-fix-syncplugin-test-driver' into 'main'
Mark Andrews [Wed, 8 Jul 2026 02:08:38 +0000 (12:08 +1000)]
Properly check test-syncplugin skip label for equality
The label length and the string length where not being checked so
a label that started with a string that matched the skip label would
incorrectly match.
Mark Andrews [Thu, 9 Jul 2026 04:47:28 +0000 (14:47 +1000)]
cdnxdomain test is failing on some platforms
Ensure the modification time is newer (second granuality) when the
zone file is rewritten as named uses the file modification time to
determine if it needs to reload a file.
Aydın Mercan [Thu, 12 Feb 2026 07:07:57 +0000 (10:07 +0300)]
clear the error stack at the end of fetching
Clearning the error stack at the very end will get rid of any other
optional fetch failures since fetch failures are treated as-if they are
unsupported by the provider.
Aydın Mercan [Thu, 12 Feb 2026 06:08:56 +0000 (09:08 +0300)]
add quic header protection to isc_crypto
QUIC uses a custom PRF construction to protect parts of the packet
header. This PRF is derived from the negotiated AEAD key and uses
unauthenticated encryption internally.
We do not expose the primitives underneath (AES-ECB and ChaCha20) as
they shouldn't be within the reach of contributors for their own safety.
Allowing such functionality to be used easily can only result it
problems not to dissimilar to leaving a baby with open bottles of
cleaning supplies.
Aydın Mercan [Mon, 9 Feb 2026 05:03:02 +0000 (08:03 +0300)]
add aead api to isc_crypto
The new AEAD API exists to cather to the needs for QUIC but is still
usable in other future contexts. Only AES-128-GCM, AES-256-GCM and
ChaCha20-Poly1305 are supported.
AES-128-CCM is intentionally skipped as the algorithm is neither
encountered in the wild nor has any useful advantages comapred to the
more popular AES-128-GCM mode.
Aydın Mercan [Tue, 3 Feb 2026 07:56:00 +0000 (10:56 +0300)]
add types compatible builtin
This builtin function makes macros gain type safety and also the ability
to statically assert the correctness of typedefs that target external
libraries.
Martin Basti [Tue, 14 Jul 2026 13:08:01 +0000 (13:08 +0000)]
fix: test: Replace python deprecated datetime utc functions
Functions `utcnow` and `utcfromtimestamp` are deprecated in python and
print warnings into tests logs about it.
Use the python prefered way by defining `timezone.utc` in `now` and
`fromtimestamp` functions. Which are equivalent but safer than naive
objects without timezone.
To ensure comptibility `%z` was added to format string to properly
process `Z` as UTC timezone.
Assisted-by: Claude Code:claude-opus-4-8[1m]
Merge branch 'mbasti/python-fix-deprecated-utcfromtimestamp' into 'main'
Martin Basti [Tue, 14 Jul 2026 09:09:10 +0000 (11:09 +0200)]
Replace python deprecated datetime utc functions
Functions `utcnow` and `utcfromtimestamp` are deprecated in python and
print warnings into tests logs about it.
Use the python prefered way by defining `timezone.utc` in `now` and
`fromtimestamp` functions. Which are equivalent but safer than naive
objects without timezone.
To ensure comptibility `%z` was added to format string to properly
process `Z` as UTC timezone.
Martin Basti [Tue, 14 Jul 2026 12:24:27 +0000 (12:24 +0000)]
new: usr: Built-in hints can be printed with named -H command
Additionally root hints were updated to precisely match authoritative source including comments. This is a cosmetic change IP addresses haven't been changed.
new: dev: Add development guidance for AI coding agents under .agents/skills/
This adds a set of skill documents that give AI coding agents the
project's established practices up front instead of having them
rediscovered (or gotten wrong) in every session: the canonical build
and test invocations, the memory-allocator contract, the disciplines
for RCU mutation, per-loop sharded structures, struct-layout work, and
flight-recorder debugging of concurrency bugs, plus the commit and
merge-request conventions.
Merge branch 'ondrej/add-agents-skills' into 'main'
Explain that MR titles and descriptions feed the generated release
notes, so agents must write them for system administrators: one short
paragraph in operational terms, no internal names or jargon, no
hand-written doc/notes/ entries, and bug framing rather than security
framing for local-filesystem misbehavior.
Walk agents through the full commit workflow: the git-clang-format
staging sequence, reason-focused messages hard-wrapped at 72 columns
with no type prefixes, the Assisted-by trailer and the forbidden ones,
amend and fixup discipline for HEAD and non-HEAD commits, and the rule
that agents commit locally and leave publishing to the user.
Capture the ownership-instead-of-locking pattern for per-loop sharded
structures: owner-only mutation under isc_tid() affinity, foreign
deletion as mark plus wait-free handoff of the exact entry to the
owner (never an O(shard) scan for marked entries), shard-held
references with bounded zombie lifetime, and eviction pressure spread
across shards instead of draining one before the next.
Point agents at pahole on the developer build's DWARF instead of
compiling throwaway sizeof programs, and document the cacheline-padding
idiom (union arm with a plain ISC_OS_CACHELINE_SIZE multiplier plus a
STATIC_ASSERT) over the enumerated-sizeof formula, which silently
miscounts when members are added.
Add the lttng-tracing-root-cause-analysis agent skill
Describe the LTTng flight-recorder methodology for concurrency bugs
that static reading, printf and debuggers all miss: small snapshot
buffers to keep timing faithful, a self-diagnosing violation tracepoint
followed by snapshot-and-abort, and the trace-reading patterns —
notably that a stale-read-after-write "paradox" indicates a missing
happens-before edge, not a timing problem.
Capture the build-invisible/publish/reclaim discipline for mutating
RCU-read structures: what makes a node observable (forward, backward
and secondary-index channels), why allocation failure must stay in the
invisible phase, publish-ordering rules for reader consistency, and the
anti-patterns (mutate-then-rollback, wiring clusters via read-side
recovery) that lead to use-after-free. Includes a worked
compressed-split example.
Condense the isc_mem/isc_mempool contract into agent guidance: the two
allocation families and why mixing them detonates the inuse INSIST at
context destroy, the pointer-NULLing put/free macros, water-mark and
striped-statistics behavior, ISC_MEM_DEBUG* facilities, and mempool
locking/ASAN caveats. Ends with a review checklist for allocation code.
fix: usr: Properly prevent TSIG generation command line injection attacks
When key names are generated with `rndc-confgen`, `tsig-keygen` and `ddns-confgen`, special characters must be escaped to ensure the configuration is parsed correctly.
Closes #6071
Merge branch '6071-allow-all-valid-keynames' into 'main'
Mark Andrews [Tue, 7 Jul 2026 01:02:34 +0000 (11:02 +1000)]
Update tests_rndc_confgen.py to show escaped double quotes
The old INJECTION string was failing due to not being a valid
DNS name providing a false assertion that injections where
no longer possible. Shorten it to fit in a single label then
check that it is properly escaped to prevent the injection attack.
Mark Andrews [Tue, 5 May 2026 01:54:35 +0000 (11:54 +1000)]
Allow all valid key names
TSIG keys names need to be able to be set to any valid name so that
update self rules can work for any valid name. Restore this ability
to the key generating tool while preventing rndc.conf and named.conf
from being compromised due to specially crafted key names.
Mark Andrews [Mon, 6 Jul 2026 02:59:50 +0000 (12:59 +1000)]
Add DNS_NAME_QUOTED flag for dns_name_totext()
Names that are to be printed within a pair of double quotes,
(for example, in named.conf), don't need spaces and special
characters to be fully escaped.
fix: usr: Ensure NSEC authority does not cross zonecut boundary
When using a cached NSEC record to prove that a delegation is insecure,
we now check that the signer name in the corresponding RRSIG is not
above a known secure delegation point. This prevents a signed namespace
from being downgraded to insecure using an NSEC record from the
grandparent zone.
Alessio Podda [Thu, 4 Jun 2026 14:40:03 +0000 (16:40 +0200)]
Add a system test for the grandparent NSEC downgrade
A resolver must not accept an NSEC or NSEC3 record signed by a zone
above a known secure delegation as proof that the delegation is
insecure. Otherwise anyone able to answer for the grandparent can
downgrade the signed namespace below it and serve forged, unsigned
records for any name in it.
Checking for SERVFAIL alone would not pin this down: a resolver that
rejects the forged proof but keeps walking down fails too, on the
unvalidatable answer it meets further along. The tests therefore
assert the refusal itself -- it is logged, and no DS query for a name
below the forged proof ever reaches the authoritative server -- and
they do so both for a proof fetched on demand and for one already in
the cache.
Evan Hunt [Fri, 22 May 2026 02:34:00 +0000 (19:34 -0700)]
Ensure NSEC authority does not cross zonecut boundary
When using a cached NSEC record to prove that a delegation is insecure,
we now check that the signer name in the corresponding RRSIG is not
above a known secure delegation point. This prevents a signed namespace
from being downgraded to insecure using an NSEC record from the
grandparent zone.
chg: dev: Pass the work callback result to the done callback
The `isc_work` callback now returns `isc_result_t` and the value is
handed to the done callback, so the callers no longer need their own
result-passing state.
Merge branch 'ondrej/pass-result-from-work-callback' into 'main'
Pass the work callback result to the done callback
The isc_work callback returned void, so every user that cared about
the outcome of the offloaded work had to smuggle it through its own
context state (xfrin_work_t, the catz/rpz updateresult fields, result
members in the dump/load/checksig contexts). Make the work callback
return isc_result_t and have isc_work deliver that value to the done
callback.
fix: usr: Prevent aborts during expired cache dumps
Running rndc dumpdb -expired could cause named to abort when the cache contained internal deletion markers for records that had already been removed. BIND now skips those markers when preparing expired cache dumps, so the dump includes only real cached records and completes normally.
Closes #6064
Merge branch '6064-skip-nonexistent-headers' into 'main'
Add a regression test that deletes a cached rdataset and then walks
all rdatasets with expired entries allowed. The iterator must report no
datasets for the deleted type rather than exposing the tombstone.
The cache rdataset iterator must never bind delete tombstones, even
when expired cache entries are requested for dumpdb. Treat non-existent
slab headers as inactive so expired dumps cannot expose headers without
a backing rdataslab.
Remove prereq.sh support from the system test runner
With every prereq.sh converted to pytest markers, the conftest
fixture no longer needs to locate and run a per-directory prereq.sh.
Drop check_prerequisites() and the README entry for the file.
The libxml2/json-c requirement becomes with_libxml2_or_json_c. The
stale Net::DNS check and the core-Perl File::Fetch check are dropped
rather than carried over.
Nicki Křížek [Tue, 30 Jun 2026 14:53:10 +0000 (14:53 +0000)]
Move Perl-module prereq.sh checks to pytest markers
fetchlimit and nsupdate still invoke ungated Perl helpers that
need Net::DNS (ditch.pl, packet.pl); reclimit and serve_stale
still run Perl ans.pl servers needing Net::DNS::Nameserver and
Time::HiRes. Replace the directory-scoped prereq.sh with new
runtime-probing markers applied only to the tests.sh wrappers
that actually run the Perl code -- native pytest modules in the
same directory (e.g. nsupdate) no longer skip when these modules
are absent.
fix: usr: Negative caching stopped working with stale-answer-client-timeout 0
With "stale-answer-client-timeout 0" configured, every client query for a
name cached as NXDOMAIN or NODATA was sent on to the authoritative servers,
even while the cached negative answer was still within its TTL, so the
resolver effectively lost negative caching. Negative answers are now
refreshed only once they have actually gone stale.
Closes #6245
Merge branch '6245-fix-query_stale_refresh_ncache' into 'main'
Test that a fresh negative cache entry is not refreshed
The existing serve-stale tests all use negative answers with a two
second TTL, because they are there to exercise stale data. Nothing
covered the far more common case of a negative answer that is still
fresh, which is how the needless refresh went unnoticed.
ans2 grows a NODATA and an NXDOMAIN name backed by a SOA with a 600
second TTL and MINIMUM, so the cached entry cannot go stale while the
test runs, and the test counts the queries that reach ans2: priming the
cache may send one, the repeated client queries must send none.
Only refresh negative cache entries that are actually stale
query_ncache() always passed a NULL rdataset to query_stale_refresh(),
which reads NULL as "this RRset is stale". NULL is only meaningful for
the DNS64 caller, whose rdataset has already been detached by the time
the answer is turned into an NXDOMAIN; everywhere else a perfectly fresh
negative cache entry was taken for a stale one.
With stale-answer-client-timeout 0 the staleness check is the only gate
left on the refresh, so every client query for a cached NXDOMAIN or
NODATA name started another fetch and negative caching stopped having
any effect.
Aydın Mercan [Fri, 3 Jul 2026 11:49:15 +0000 (14:49 +0300)]
[CVE-2026-13321] sec: usr: Fix DNSSEC validation bypass via out-of-zone NSEC Next Field
A malicious zone with out-of-zone NSEC next owner names can cause a DNSSEC validating resolver to cache such record and, if `synth-from-dnssec` is enabled, to generate negative answers for any zone that is covered by the range.
ISC would like to thank Qifan Zhang of Palo Alto Networks for reporting the issue.
Closes isc-projects/bind9#5873
Merge branch '5873-security-out-of-zone-nsec-dnssec-bypass' into 'security-main'
Evan Hunt [Thu, 14 May 2026 03:45:57 +0000 (20:45 -0700)]
dns_rdataset_addnoqname() could find unsigned NSEC/NSEC3
The dns_rdatalist addnoqname() implementation searches for the first
NSEC or NSEC3 record in a message, then for the first RRSIG covering
that type in the same message. Previously, if no RRSIG for the type was
found, the function accepted the unsigned record. Now, it will instead
continue searching until an NSEC or NSEC3 that does have a matching
signature is found.
When this function is called from validated() in resolver.c, a
non-success return code is now treated as an error instead of triggering
an assertion failure.
[CVE-2026-10723] sec: usr: Correct verification of NSEC3 signer name
BIND 9 accepted child-zone NSEC3 records where the first label equals the hash of the parent zone as valid parent-zone closest encloser proofs. This has been fixed.
ISC thanks Qifan Zhang of Palo Alto Networks for reporting the issue.
Closes isc-projects/bind9#5874
Merge branch '5874-confidential-nsec3-apex-hash-bypass' into 'security-main'
Update the llm generated reproducer:
- Move server.py into ans1/ans.py
- Remove unnecessary named.conf configuration options
- Add comments describing the steps (copied from GL issue)
- Rename system test
Aydın Mercan [Thu, 7 May 2026 15:59:20 +0000 (18:59 +0300)]
Reject out-of-zone NSEC next owner names
When verifying DNSSEC records, make sure that a next owner name of
an NSEC record is a subdomain of the signer field.
This follows the specification RFC 4034, section 4.1.1:
Owner names of RRsets for which the given zone is not authoritative
(such as glue records) MUST NOT be listed in the Next Domain Name
unless at least one authoritative RRset exists at the same owner
name.
While the above paragraph is intended for glue records, it also
applies to out-of-zone data.
sec: usr: Reclaim memory promptly when DNSSEC validations are canceled
When a resolver is flooded with queries that require DNSSEC
validation — for example during a random-subdomain attack —
many of those validations are canceled before they complete.
Previously a canceled validation still kept its place in the
internal work queue and held the associated response in memory
until that queued work eventually ran, so memory could climb
sharply under sustained load. The canceled work is now dropped
as soon as the validation is canceled, releasing the memory it
was holding.
Merge branch '4760-cancel-dns_validator-jobs-early' into 'security-main'
Ondřej Surý [Tue, 23 Jun 2026 05:03:43 +0000 (07:03 +0200)]
Make the per-fetch validation quota terminal in sub-validations
Negative-proof validation stops once the per-fetch validation quota is
exhausted, but the DNSKEY, DS and CNAME sub-validator callbacks did
not — they re-fetched the record or relabelled the quota as a broken
chain. Treat ISC_R_QUOTA as terminal in all three, as
validator_callback_nsec already does.
Aydın Mercan [Wed, 6 May 2026 13:54:57 +0000 (16:54 +0300)]
Add system test for out-of-zone nsec dnssec bypass
A malicious zone with out-of-zone NSEC entries can get a DNSSEC
validating resolver's cache to cover the victim zone for non-existence
and prevent nameserver queries without DNSSEC failure.
Test for this case with an `evil.test` zone that tries to cover the
`victim.test` zone.
Colin Vidal [Wed, 1 Jul 2026 09:56:41 +0000 (11:56 +0200)]
chg: dev: Follow-up of disambiguate `query_cname()` and `query_dname()` usage
Previous commit "Disambiguate `query_cname()` and `query_dname()` usage"
was harmless but also useless, as it was checking `qctx->result` which
is always set to `ISC_R_SUCCESS` when `qctx` is initialized. The intent
was to check `qctx->fresp->result` (which is the result provided by the
resolver). But this was also wrong (this is actually the case we do
expect `query_cname()`/`query_dname()` to be called, to follow the
chain).
The actual invarant that needs to be checked is if the qtype is CNAME
then we do not follow the chain, so we can't call `query_cname()`. This
invariant has been added.
If the qtype is DNAME, it's more complex, because a DNAME can be found
from a local zone or cache and the chain can be locally followed. In
which case, calling `query_dname()` is legit, as soon as the qname is a
subname of the DNAME target. This invariant is already checked.
Merge branch 'colin/follow-up-disambiguate-query_cname_dname' into 'security-main'
Ondřej Surý [Fri, 12 Jun 2026 15:37:16 +0000 (17:37 +0200)]
Cancel the offloaded verification job when canceling a validator
The offloaded jobs were fire-and-forget: dns_validator_cancel() could
only raise a flag and wait for the queued crypto to get its turn, so
under a random-subdomain attack canceled validations piled up in the
worker queues, pinning memory. Keep the isc_work handle and cancel it:
a still-queued job never runs its crypto and unwinds on dequeue; a
running one is unaffected.
Colin Vidal [Wed, 1 Jul 2026 08:59:01 +0000 (10:59 +0200)]
Convert some chain shell-based system tests in python
Rewriting a couple of shell-based chain system test into python. Those
tests exercise the dname resolution of either an authoritative or
resolver without extra queries (it is self resolver either because the
target is in an authoritative zone or in the same authoritative answer).
Also add an extra test checking the same mechanism, this time, from the
resolver cache (no resolver involed).
Those highlight the fact that `query_dname()` can be called in situation
where the query type is DNAME, and even though, it needs to be resolved.
Colin Vidal [Wed, 24 Jun 2026 06:40:52 +0000 (08:40 +0200)]
chg: dev: Disambiguate `query_cname()` and `query_dname()` usage
Make explicit the fact that `query_cname()` and `query_dname()` must be
called only from a context where the resolver is answering a question
which is _not_ respectively `CNAME` or `DNAME`.
Merge branch 'colin/explicit-dname-cname-libns-query-usage' into 'security-main'
Colin Vidal [Wed, 1 Jul 2026 08:58:41 +0000 (10:58 +0200)]
Follow-up of disambiguate `query_cname()` and `query_dname()` usage
Previous commit "Disambiguate `query_cname()` and `query_dname()` usage"
was harmless but also useless, as it was checking `qctx->result` which
is always set to `ISC_R_SUCCESS` when `qctx` is initialized. The intent
was to check `qctx->fresp->result` (which is the result provided by the
resolver). But this was also wrong (this is actually the case we do
expect `query_cname()`/`query_dname()` to be called, to follow the
chain).
The actual invarant that needs to be checked is if the qtype is CNAME
then we do not follow the chain, so we can't call `query_cname()`. This
invariant has been added.
If the qtype is DNAME, it's more complex, because a DNAME can be found
from a local zone or cache and the chain can be locally followed. In
which case, calling `query_dname()` is legit, as soon as the qname is a
subname of the DNAME target. This invariant is already checked.
Ondřej Surý [Tue, 23 Jun 2026 05:05:01 +0000 (07:05 +0200)]
[CVE-2026-11605] sec: usr: Prevent excessive validation work from crafted negative responses
A validating resolver could be made to perform a large amount of DNSSEC
validation work in response to a single answer, consuming excessive CPU. A
malicious authoritative server triggers this by returning a signed negative
answer (NXDOMAIN or NODATA) padded with many denial-of-existence proof
records, which the resolver continued to verify beyond its per-query
validation limit. It now enforces that limit on negative answers and returns
SERVFAIL once the limit is reached.
Closes: https://gitlab.isc.org/isc-projects/bind9/-/work_items/4463
Merge branch '4463-limit-the-number-of-negative-validations' into 'security-main'