Before you begin
Networking foundations and configuration labs →
Start with the Networking foundations course. Use an isolated lab for configuration, traffic generation and failure injection. Production examples are read-only unless a separately authorized change plan says otherwise. Commands are platform examples, not universal syntax; record model and software before using them.
Working toward: Trace a service end to end, verify each dependency, diagnose faults with evidence, measure recovery and leave a reproducible operational record.
Read each explanation, run the example in your own lab, and attempt the exercise before opening its answer. Published lessons are ready to study; unfinished roadmap topics remain planned.
Validation: Address-plan arithmetic and content structure are checked. Commands and scenarios are documented examples, not claims of execution on physical routers, firewalls, radios or carrier circuits. Actual acceptance requires the matching environment and measured results.
1. Translate requirements into a testable design
Start with services and traffic, then choose topology. For each service record the client and server networks, direction, transport, destination ports, name-resolution dependencies, authentication, required latency/loss/jitter, expected concurrent users and transfer sizes. Separate an application’s documented requirement from an observed flow. An outbound connection can still require inbound return traffic; a stateful policy is not the same as a broad inbound permit.
Draw physical, Layer 2, Layer 3 and security views separately. The physical view shows equipment, patching, power and carrier boundaries. The Layer 2 view shows VLAN carriage, loop prevention and link aggregation. The Layer 3 view shows subnets, gateways, VRFs and routing adjacencies. The security view shows allowed flows, NAT, tunnels and management paths. Label ownership at every boundary so a fault can be assigned with evidence.
Define acceptance before making changes. Each requirement needs a test, source, destination, expected result, evidence format, owner and exception process. Include negative tests: a guest subnet must not reach management; a failed uplink must not strand the administrative path. Set numerical targets from the actual service requirement or contract rather than inventing universal thresholds.
Flow record (example, documentation addresses):
Client: 192.0.2.0/26, user VRF
Server: 198.51.100.20, application VRF
Protocol/port: TCP/443
Dependencies: DNS, trusted time, identity provider
Positive test: authenticated transaction succeeds
Negative test: guest VRF cannot initiate this flow
Evidence: UTC time, endpoints, policy/session ID, resultWhat to expect
A requirement-to-test matrix, four consistent diagrams and named owners for unresolved dependencies.
Your turn
“The network must work” is the only requirement supplied. What do you ask for before choosing bandwidth?
Show answer and reasoning
Application types, traffic direction, concurrency, peak transfer volume, response-time targets, failover behavior, available carrier services and growth. Record assumptions and resolve them before treating the design as accepted.Watch for: A link-up indication is a physical observation, not an application acceptance result.
Link to this lesson2. Survey the room, rack and path
Inventory the room and each rack: rack units, depth, rail compatibility, airflow direction, clearances, grounding/bonding arrangements, power circuits, PDU connector types, UPS runtime under measured load, environmental monitoring and physical access. Dual power supplies do not create independent power if both terminate on one PDU or upstream circuit. Trace the dependency, not just the cable count.
Record the carrier demarcation, equipment ownership, patch-panel ports, cable route, fiber type, connector polish, optic wavelength, reach and any cross-connect order. A connector fitting is not proof of optical compatibility. UPC and APC connectors should not be mixed. Keep bend radius and cleanliness within the component requirements; inspect and clean fiber with appropriate equipment, never by looking into an active connector.
Create a port schedule before cabling. Include device, slot/port, local patch-panel port, remote port, media, optic part number, VLAN or routed role, speed, expected negotiation, label and test result. Photograph labels and rack placement where permitted. Keep spare supported optics, patch leads, console adapters and a tested administrative laptop available.
Port schedule:
SW-A Gi1/0/48 → panel A-12 → panel B-12 → SW-B Gi1/0/48
Role: trunk; VLANs 10,20,99; LACP member group 1
Media: exact fiber/optic pair recorded separately
Power: PSU1→PDU-A→circuit A; PSU2→PDU-B→circuit BWhat to expect
Every installed path can be traced in both directions, with matching media and genuinely understood failure domains.
Your turn
Two uplinks fail whenever one patch-panel enclosure is disturbed. Why did the redundant design fail?
Show answer and reasoning
The links share a physical risk. Separate cable routes and termination points where the availability requirement demands it; validate upstream equipment and power diversity too.Watch for: Electrical work, grounding changes and mains measurements require the appropriate qualified personnel, not an improvised network workaround.
Link to this lesson3. Stage equipment and preserve a recovery path
Record model, serial number, modules, software image, boot variables, license state, support entitlement and hardware compatibility. Verify the proposed software release against the actual platform and feature requirements. A command accepted by a simulator does not prove that a hardware model supports the feature or scale.
Build a minimal baseline: hostname, management address/VRF, explicit administrative access, trusted time, logging, secure management protocols and a tested local recovery method. Stage application-facing configuration separately so the administrative path can be verified first. Back up the original running and startup configurations, software and relevant certificates securely; redact secrets from ordinary reports.
Test console access with the exact cable and terminal settings before depending on it. Confirm out-of-band access does not traverse the network you are about to change. For platforms supporting transactional configuration, validate a candidate and use a confirmed commit with a realistic timer. Otherwise prepare a platform-specific rollback and an onsite/console operator. Saving the running configuration is not the same as proving a reboot will recover it.
Cisco IOS-like read-only inventory (platform-dependent):
show version
show inventory
show boot
show license summary
show running-config
show startup-config
Junos example of a guarded change workflow, isolated lab:
configure
# apply reviewed candidate
commit check
commit confirmed 10
# verify through an independent session
commitWhat to expect
Known software/hardware baseline, tested independent access, usable backup and an executable rollback procedure.
Your turn
Your change replaces the only management default route. What must exist before applying it?
Show answer and reasoning
A verified console or independent management path, the exact old route/configuration, rollback authority and a bounded verification window. A second SSH session over the same route is not independent.Watch for: Example commands are not vendor-neutral; confirm syntax and feature behavior for the exact release.
Link to this lesson4. Run a bounded change and rollback
A change plan states scope, owner, dependencies, start/stop time, affected services, exact actions, verification and rollback. Put irreversible or slow operations early enough to recover before the window ends. Record a latest safe rollback time based on the restoration duration, not merely the scheduled end.
Use explicit checkpoints: baseline captured; management reachable; physical links healthy; Layer 2 stable; routing correct; policy correct; services functional; failover verified; monitoring operational. At a failed checkpoint, freeze unrelated changes, gather evidence and decide whether the bounded diagnosis time has expired. Changing several layers simultaneously destroys causal evidence.
Rollback means restoring a known service state, not merely undoing the last command. New DHCP leases, NAT/session state, routing advertisements, DNS caches and endpoint configuration may survive a configuration reversal. Include these dependencies in the recovery procedure and retest the same service transactions used for acceptance.
Checkpoint log:
UTC | step | operator | observation | evidence | decision
Stop condition: management lost, route leak, loop, or agreed error threshold exceeded
Latest rollback start: window end minus measured restore time and verification buffer
Recovery validation: same application tests as pre-change baselineWhat to expect
A second engineer can execute the plan and know precisely when to stop or restore.
Your turn
Configuration rollback succeeds, but users still resolve the new address. What was omitted?
Show answer and reasoning
DNS TTL/cache behavior and possibly client/application caching. Verify authoritative and recursive answers and the client’s actual resolver; avoid indiscriminate cache clearing before capturing evidence.Watch for: Do not confuse permission to diagnose with permission to change every neighboring system.
Link to this lesson5. Understand the carrier handoff
Obtain circuit/service identifiers, demarcation location, support contacts, handoff medium/speed, encapsulation, VLAN tags, committed information rate, burst parameters, MTU definition, addressing, routing arrangement and contractual test criteria. Distinguish port speed from purchased throughput. A 1 Gb/s port can carry a service policed at 100 Mb/s.
Clarify whether the service is Internet transit, point-to-point Ethernet, multipoint Ethernet, MPLS L3VPN, broadband with CGNAT or an overlay underlay. An Ethernet handoff does not tell you whether the carrier is transporting your VLAN transparently or terminating Layer 3. Record CE/PE responsibilities, routing policy and which party owns the next hop.
For tagged handoffs confirm customer tag, provider tag if relevant, expected untagged/native behavior and priority markings. For routed handoffs confirm address/mask, gateway, ASN/neighbors if used, static routes and allowed prefixes. Ask the provider for a loopback or test boundary only through the approved procedure; a loopback can interrupt service.
Carrier evidence bundle:
Circuit ID and demarc port
UTC fault start and recurrence
Local admin/oper state, speed, optic levels and counter deltas
Expected vs observed VLAN/encapsulation and next-hop ARP/ND
Tests to demarc, provider next hop and agreed remote endpoint
No secrets or unrestricted full packet captureWhat to expect
Both parties agree on exactly what crosses the demarc and what the provider is expected to deliver.
Your turn
The physical port is 1 Gb/s but sustained traffic drops above 100 Mb/s. What hypotheses fit?
Show answer and reasoning
A 100 Mb/s policer/CIR, shaping, a downstream bottleneck or an endpoint limit. Compare contract, policer counters, queue drops and controlled tests before declaring a faulty optic.Watch for: A successful provider loopback test does not prove customer routing, DNS or application policy.
Link to this lesson6. Diagnose copper, optics and negotiation
Separate administrative down, no signal, negotiation failure, link flaps, errors and congestion. Inspect both ends. Record counter deltas over a known interval instead of treating lifetime CRC counts as current faults. CRC/FCS errors suggest corrupted frames; output drops commonly indicate queue pressure; they are not interchangeable.
For copper check category, certified length, pairs, patching, port capability, speed and duplex. For fiber compare receive/transmit power with the optic’s own thresholds and link budget; a received level can be too high as well as too low. Check fiber mode, wavelength, lane/polarity mapping, supported transceiver and far-end transmit state. At higher speeds verify FEC and breakout compatibility, not just the nominal speed.
For PoE record the powered device’s requested class, delivered power, switch per-port and total budgets and LLDP negotiation. A device can power on yet disable radios or reboot under load. Restore reliable physical transport and power before changing VLAN or routing policy.
Cisco IOS-like, confirm platform support:
show interfaces status
show interfaces counters errors
show interfaces transceiver detail
show power inline
show logging
Compare counters at T0 and T1; retain elapsed seconds and offered load.What to expect
Stable link state, matching negotiation, optic levels within device limits and no unexplained growing physical-error counters.
Your turn
CRC counters stay flat but output discards rise during backups. What should you inspect next?
Show answer and reasoning
Queue occupancy, offered rate, oversubscription, shaping/policing, pause frames and endpoint behavior. This is stronger evidence for congestion than corrupted media.Watch for: Forcing speed/duplex or disabling FEC without understanding both ends can create a different fault.
Link to this lesson7. Build an address and VLAN plan
Allocate addresses by required host count, growth, summarization and failure/security domains. Record subnet, mask/prefix, gateway, DHCP pool, exclusions, static reservations, DNS/NTP, VRF, VLAN and owner. VLAN identifiers and IP prefixes are separate namespaces; VLAN 20 does not imply any particular subnet.
Use VLSM from larger to smaller requirements. For ordinary IPv4 subnets, network and directed-broadcast addresses are not host addresses; /31 point-to-point links and /32 routes have different semantics. Avoid overlapping allocations within one routing context. Overlap in separate VRFs still complicates NAT, shared services and later integration.
Plan management, transit, infrastructure and users deliberately. Keep unused ranges documented, not silently consumed. For IPv6, assign /64 to normal SLAAC LANs and use a consistent site/subnet allocation; point-to-point and loopback exceptions must follow the actual design. Documentation-only addresses below must never be mistaken for production allocations.
Worked allocation from 192.0.2.0/24:
Users: 192.0.2.0/26 hosts .1–.62 broadcast .63
Voice: 192.0.2.64/27 hosts .65–.94 broadcast .95
Management: 192.0.2.96/28 hosts .97–.110 broadcast .111
Transit: 192.0.2.112/30 hosts .113–.114 broadcast .115
Reserve: remaining space, allocated only through IPAM
IPv6 example LAN: 2001:db8:20:10::/64What to expect
No overlaps, correct usable ranges, enough capacity and matching gateway/DHCP/route records.
Your turn
Allocate a subnet for 50 ordinary IPv4 hosts and explain why /27 fails.
Show answer and reasoning
/26 has 64 addresses and normally 62 usable hosts. /27 has 32 addresses and normally 30 usable hosts.Watch for: A spreadsheet that lists only host IPs but omits masks and VRFs cannot establish uniqueness.
Link to this lesson8. Verify access ports, trunks and MAC learning
An access port classifies untagged traffic into a VLAN. A trunk carries multiple VLANs, normally using 802.1Q tags with any untagged/native behavior explicitly agreed. Confirm allowed VLANs, native VLAN, endpoint expectations and voice VLAN behavior at both ends. A trunk can be up while silently excluding the only VLAN that matters.
Trace one client MAC from access port through uplinks. Verify its VLAN, aging and whether it moves unexpectedly. Then inspect the gateway’s ARP/ND entry. A learned MAC proves Layer 2 observation at that point, not that the gateway has the correct IP or route. Unknown-unicast flooding, broadcast and multicast are different forwarding cases.
Keep VLAN creation, trunk permission and Layer 3 gateway configuration as separate checklist items. Native VLAN mismatches can produce connectivity leaks or control-plane confusion. Do not solve an unknown trunk fault by allowing every VLAN permanently; test the intended set.
Cisco IOS-like:
show vlan brief
show interfaces trunk
show interfaces Gi1/0/10 switchport
show mac address-table address 0011.2233.4455
show ip arp 192.0.2.10
show interfaces Gi1/0/10
# Port names and syntax vary.What to expect
The expected client MAC appears in the expected VLAN at each hop, and the intended gateway resolves correctly.
Your turn
The client MAC appears on the access switch but not on distribution. Where do you narrow the fault?
Show answer and reasoning
The intervening trunk/VLAN path, STP forwarding state, LACP membership and physical link. Compare allowed/active VLANs and counters on both ends before editing the client.Watch for: A switch MAC table and a host ARP cache answer different questions.
Link to this lesson9. Prevent loops and verify aggregation
Select the spanning-tree mode and intended root for each topology or instance. Map blocked/alternate paths and verify the root bridge, priorities and cost decisions. Edge/PortFast behavior belongs on genuine edge ports; BPDU guard can stop an unexpected bridge there. Root guard, loop guard and unidirectional-link detection address different failure modes and require platform-specific placement.
An LACP bundle is one logical link only when both sides agree. Match speed, trunk/access role, VLAN set, native VLAN, MTU and aggregation parameters. Verify each member is collecting/distributing; a configured channel number alone proves nothing. Cross-chassis aggregation requires a supported stack/MLAG design with its own peer-link and split-brain behavior.
A single flow usually hashes to one member; two 1 Gb/s links do not imply a single TCP flow reaches 2 Gb/s. Test multiple flows when measuring aggregate capacity. Remove or fail one member in an isolated lab and verify forwarding without a loop or unintended outage.
Cisco IOS-like:
show spanning-tree root
show spanning-tree inconsistentports
show etherchannel summary
show lacp neighbor
show interfaces port-channel 1
show logging
# Read platform-specific meanings of bundled/suspended/individual flags.What to expect
Intended STP roots and roles, consistent bundle membership, and measured member-failure behavior.
Your turn
One LACP member is suspended after a trunk change. What should be compared?
Show answer and reasoning
Both physical members and both peers: VLAN/native settings, mode, speed/duplex, MTU, LACP actor/partner state and platform consistency logs.Watch for: Disabling STP to make a blocked link forward can create a broadcast storm.
Link to this lesson10. Verify gateways, ARP/ND and first-hop redundancy
A host sends off-subnet traffic to a gateway’s MAC, not directly to the remote server’s MAC. Check the host’s address/mask, selected route, gateway neighbor entry and the gateway interface state. A wrong mask can make a remote address appear local, causing unanswered ARP rather than routed traffic.
HSRP/VRRP and similar protocols provide a virtual gateway address/MAC with active/backup behavior. Record priorities, preemption policy, timers, tracking objects and expected active device. Tracking only the LAN interface may leave a gateway active after its WAN route has failed. Redundant gateways also need reachable return paths and correct routing convergence.
Duplicate addresses, stale neighbor entries, proxy ARP and security filters can make symptoms misleading. Capture before clearing. Observe gratuitous ARP or unsolicited neighbor advertisements during a failover and confirm clients update the forwarding destination.
Linux host:
ip -brief address
ip route get 198.51.100.20
ip neigh show
Cisco IOS-like gateway:
show ip interface brief
show ip arp
show standby brief
show vrrp brief
show trackWhat to expect
A correct virtual gateway, stable neighbor mapping and successful traffic through the intended active path and its backup.
Your turn
The active gateway loses upstream reachability but its LAN stays up. Why may failover not occur?
Show answer and reasoning
The redundancy protocol may not track the actual upstream dependency. Review route/interface/IP-SLA tracking and the design’s intended failover conditions.Watch for: Clearing all neighbor tables can temporarily hide duplicate-address evidence.
Link to this lesson11. Operate IPv6 as its own working path
IPv6 uses Neighbor Discovery and Router Advertisements rather than IPv4 ARP. A link-local address is valid only on its link and may require an interface scope identifier. RA supplies default-router information; DHCPv6 does not by itself replace that function. SLAAC uses advertised prefix information, and clients can have stable and temporary addresses simultaneously.
Check RA flags, prefix lifetimes, DNS delivery, DHCPv6/relay where used, duplicate-address detection and neighbor state. A global address alone does not prove a default route. Apply IPv6 firewall rules intentionally; a secure IPv4 policy does not automatically secure IPv6.
ICMPv6 is required for important functions, including Packet Too Big messages. IPv6 routers do not fragment transit packets; path-MTU failures can therefore break large transfers while small pings work. Test AAAA resolution and the actual IPv6 application path independently of IPv4, especially when clients prefer IPv6.
Linux:
ip -6 address
ip -6 route
ip -6 neigh
ping -6 -c 3 2001:db8:20:10::1
# Link-local example needs the real interface: fe80::1%eth0
Windows PowerShell:
Get-NetIPConfiguration
Get-NetRoute -AddressFamily IPv6
Resolve-DnsName example.net -Type AAAAWhat to expect
Correct addresses, route, DNS, neighbor state, security policy and successful application transactions over IPv6.
Your turn
A client receives DHCPv6 addressing but has no default route. Which control traffic do you inspect?
Show answer and reasoning
Router Advertisements, router lifetime and interface/VLAN delivery. DHCPv6 addressing alone is not proof of an IPv6 default router.Watch for: Dropping all ICMPv6 breaks essential network behavior.
Lesson references
Link to this lesson12. Follow the actual forwarding decision
The forwarding table uses longest-prefix match. A more specific route wins over a default regardless of the default’s low administrative distance. Administrative preference selects among candidate routes to the same prefix; a routing protocol’s metric compares paths according to that protocol. Policy-based routing, VRFs and source-specific rules can modify the path a simple global-table lookup suggests.
Check route source, next-hop reachability, outgoing interface and recursive resolution. The routing information base and programmed forwarding table can diverge during faults or scale issues. Verify the reverse path separately; asymmetric routing is not automatically wrong, but a stateful firewall or strict reverse-path check may reject it.
Use source-specific ping/traceroute from the correct VRF/interface. A router ping sourced from a management loopback can fail while user traffic succeeds, or the reverse. Traceroute hop loss can reflect control-plane rate limiting; correlate end-to-end loss and captures before blaming an intermediate router.
Cisco IOS-like:
show ip route 198.51.100.20
show ip cef 198.51.100.20 detail
show ip route vrf USERS 198.51.100.20
show ip policy
Linux:
ip rule show
ip route show table all
ip route get 198.51.100.20 from 192.0.2.10What to expect
The exact source/VRF/destination path is explained in both directions, including recursive next hops and policy.
Your turn
A /24 via the wrong router coexists with a correct default route. Which is used for an address inside the /24?
Show answer and reasoning
The /24, because longest-prefix match precedes comparison with the less-specific default.Watch for: A routing-table screenshot without source context and return-path evidence is incomplete.
Link to this lesson13. Bring OSPF neighbors to a useful state
Compare area, subnet, network type, timers, authentication, router IDs, passive-interface settings and MTU. An adjacency stuck in ExStart/Exchange can indicate MTU or database-exchange problems; Init can mean one-way hello visibility; Down can mean no hellos or a disabled/passive path. On broadcast networks, two non-DR/BDR routers can remain 2-Way normally.
Full adjacency only proves database exchange, not that the intended prefixes are advertised or selected. Inspect the link-state database, route type, cost and forwarding table. Summarization and filtering have specific placement rules; a distribute-list may affect local route installation without doing what you expect to the shared database.
Test a path failure and recovery, measuring lost transactions and convergence. BFD can detect failure faster but is not a substitute for correct routing policy; overly aggressive timers can cause instability on overloaded devices. Record the expected DR/BDR roles and whether a restart changes them.
Cisco IOS-like:
show ip ospf neighbor
show ip ospf interface brief
show ip ospf interface
show ip ospf database
show ip route ospf
show loggingWhat to expect
Expected neighbor states, correct prefixes and measured forwarding recovery after a controlled failure.
Your turn
Two DROther routers show 2-Way on a shared Ethernet. Is that necessarily a fault?
Show answer and reasoning
No. Full adjacencies normally form with DR/BDR on that network type. Check the role relationships and intended routes.Watch for: Changing timers on only one side can break an otherwise healthy adjacency.
Lesson references
Link to this lesson14. Verify BGP sessions and route policy
A BGP TCP session and Established state are only the start. Confirm peer ASN, source address, reachability, authentication, address family and any eBGP multihop requirement. Then inspect received, accepted, selected and advertised prefixes separately. A session can be healthy while policy rejects every useful route—or exports far too much.
Apply explicit inbound and outbound prefix policy, maximum-prefix protection and intended default-route handling. Record local preference, AS path, MED, communities and next-hop behavior where relevant. BGP best-path selection is not simply shortest AS path in every case; platform policy and other attributes precede or follow it. Longest-prefix forwarding still applies after routes are selected.
At an enterprise/provider boundary, confirm whether you receive a default, selected private routes or a full table, and whether the hardware can support the scale. Test withdrawal and restoration of only the authorized prefixes. Coordinate provider changes; never announce an address block merely because it appears in a lab example.
Cisco-like, exact syntax varies:
show bgp ipv4 unicast summary
show bgp ipv4 unicast 198.51.100.0/24
show bgp ipv4 unicast neighbors 192.0.2.114 advertised-routes
show ip prefix-list
show route-map
# Received-routes visibility depends on platform and retained inbound policy state.What to expect
Only approved prefixes are accepted/exported, next hops resolve, and failover follows the intended policy.
Your turn
BGP is Established but no route is installed. Name four checks.
Show answer and reasoning
Address-family activation, inbound policy, next-hop reachability, and whether another route/path is preferred; also check advertised prefixes and maximum-prefix/session logs.Watch for: A broad permit in an outbound route policy can create a route leak.
Lesson references
Link to this lesson15. Keep routing domains and shared services explicit
A VRF separates routing/forwarding tables; it does not automatically impose application security. Record interface membership, connected prefixes, defaults and route-leak policy. A packet tested in the global table may behave differently from the same destination in a user VRF.
Shared DNS, NTP, identity and management services need deliberate bidirectional reachability. Route import/export or firewall transit must be limited to intended prefixes and flows. Overlapping addresses require NAT or another explicit design; a route leak cannot make two identical prefixes unambiguous in one table.
MPLS L3VPN uses provider control-plane mechanisms to carry separated customer routes; a customer engineer still validates CE routing and the service boundary. VXLAN/EVPN adds an underlay/overlay distinction: VTEP reachability, tunnel MTU, VNI/VLAN mapping and EVPN route state can fail independently. Learn which layer your device owns before changing a provider or fabric control plane.
Cisco-like:
show vrf
show ip route vrf USERS
show ip interface brief vrf USERS
# Platform-specific VRF ping with an explicit source
Overlay evidence:
underlay route → VTEP reachability → VNI mapping → learned MAC/IP → policyWhat to expect
Separation is preserved, shared-service flows are intentional, and tests use the correct routing domain.
Your turn
The global table reaches DNS but clients in USERS cannot. What does the global ping prove?
Show answer and reasoning
Only the global/source context used by that ping. Inspect USERS routes, source selection, route leaking, firewall policy and return routes.Watch for: VRF separation alone is not an authorization policy.
Lesson references
Link to this lesson16. Read firewall and ACL behavior from a flow
Write the five-tuple and direction before reading policy: source IP/port, destination IP/port and protocol, plus interface/zone/VRF and time. A stateless ACL evaluates packets independently; a stateful firewall tracks sessions and may permit return packets without a separate broad reverse rule. Object groups, inherited policy and rule order can change what the visible rule means.
Confirm whether policy matches pre-NAT or post-NAT addresses on the specific platform. Use session tables, rule hit counters and narrowly scoped logs. A packet capture at ingress without a corresponding egress packet narrows the location, but you still need the drop reason: policy, routing, zone mismatch, reverse-path check, inspection or resource exhaustion.
Test both approved and forbidden flows. Avoid replacing a specific rule with any-any as a permanent diagnosis. When a temporary test exception is explicitly authorized, bound its endpoints, ports, duration and removal, then preserve the observed difference. Application identification and TLS inspection can impose checks after the TCP handshake.
Flow worksheet:
192.0.2.10:ephemeral → 198.51.100.20:443 TCP
Ingress zone/VRF: USERS
Expected egress: APP
NAT before/after: documented for this platform
Rule ID, hit-count delta, session state, drop reason
Reverse flow: established return through same state domainWhat to expect
Intended positive and negative results with rule/session evidence, not merely a green ping.
Your turn
TCP establishes, then the firewall resets the flow. Which layers remain plausible?
Show answer and reasoning
Application identification, TLS inspection, policy reclassification, threat inspection, server behavior or session/resource issues. Basic reachability already worked at that instant.Watch for: Logs can contain sensitive addresses or payloads; keep evidence scoped and access-controlled.
Link to this lesson17. Trace translation without losing the return path
Distinguish source NAT/PAT, destination NAT, static one-to-one mapping and twice NAT. Record original and translated tuples in each direction. PAT shares an address by allocating ports; exhaustion can cause intermittent failures under high concurrency even when bandwidth is low.
A destination NAT rule does not create a listening server, a route, an allow policy or a valid TLS certificate. The server’s return traffic must traverse the translation state or use a compatible symmetric design. Hairpin access from inside to an external name may need a specific NAT/policy arrangement or split DNS; choose the intended architecture rather than layering unexplained translations.
Check session lifetimes for long-idle applications and UDP mappings. Carrier-grade NAT can prevent inbound initiation even when the local firewall is open. IPv6 normally does not require address translation for ordinary global reachability; policy must still be explicit.
Translation record:
Original: 192.0.2.10:51000 → 198.51.100.20:443
Translated: 203.0.113.10:62000 → 198.51.100.20:443
Return: 198.51.100.20:443 → 203.0.113.10:62000
Restored: 198.51.100.20:443 → 192.0.2.10:51000
Capture both sides using synchronized clocks.What to expect
Every observed tuple maps to a known translation and return state, with capacity and timeout behavior understood.
Your turn
Only new connections fail during peak usage; existing sessions remain healthy. What NAT evidence matters?
Show answer and reasoning
Available translation ports, allocation failures, per-source limits, session table occupancy and timeout churn, alongside firewall/endpoint limits.Watch for: Opening a destination port cannot fix a server bound only to loopback.
Link to this lesson18. Diagnose IPsec and overlay tunnels
Separate underlay reachability, IKE negotiation, child/IPsec SAs, traffic selectors, routing and application traffic. IKE success does not prove the correct protected subnets are selected. Compare identities, authentication, proposals, lifetimes, PFS/DH groups, NAT traversal and selector definitions on both peers using the exact platform terminology.
For a route-based VPN inspect the tunnel interface, routes and policy; for policy-based VPN inspect the encryption domain/crypto ACL and its interaction with NAT. NAT exemption order is platform-specific. Check encapsulated and decapsulated packet counters in both directions. Increasing encrypt counters without decrypt counters suggests a one-way problem but does not locate it without peer evidence.
Account for tunnel overhead in MTU and MSS. Test a small transaction and a large transfer. Rekey failures, idle timeout, path changes, fragmentation and clock/certificate problems can produce faults that a one-time ping misses. Collect peer logs around the same UTC interval.
Cisco IOS-like IPsec examples:
show crypto ikev2 sa
show crypto ipsec sa
show interfaces tunnel 0
show ip route 198.51.100.0
# Match identities/selectors and packet-counter deltas on both peers.
# UDP 500/4500 and ESP handling depend on NAT-T and design.What to expect
Negotiated SAs, matching selectors, bidirectional packet counters, correct routes and stable application traffic through rekey/failover.
Your turn
Tunnel status is up, encrypt counters rise, decrypt counters stay flat. What next?
Show answer and reasoning
Check remote decapsulation/return routing, remote selectors/policy/NAT, underlay filtering and whether replies enter the tunnel. Correlate captures on both peers.Watch for: Do not publish pre-shared keys, private keys or full unredacted VPN configuration.
Link to this lesson19. Find MTU, MSS and fragmentation failures
MTU limits packet size at a link; path MTU is the smallest usable value along a path. Encapsulation consumes space, so a tunnel over a 1500-byte underlay may require a smaller inner MTU or an underlay supporting larger frames. Ethernet frame size, IP packet size and test-tool payload size are different quantities.
For IPv4, a DF packet that is too large should trigger an ICMP fragmentation-needed response. For IPv6, routers send Packet Too Big rather than fragmenting transit traffic. Blocking these messages can create a black hole: small requests succeed while TLS, file transfer or certain pages stall. TCP MSS is the data payload advertised during handshake; clamping can help a known TCP path but does not fix every UDP or non-TCP MTU problem.
Test from the actual source and routing domain, vary packet size carefully and capture ICMP plus retransmissions. A failed large ping is not proof of an MTU fault if the destination blocks all echo requests. Compare a working small probe, a controlled large probe and a real application.
IPv4, 1500-byte IP packet examples (no IP options):
Linux: ping -c 3 -M do -s 1472 198.51.100.20
Windows: ping -n 3 -f -l 1472 198.51.100.20
# 1472 payload + 8 ICMP + 20 IPv4 = 1500
# Use an authorized responding endpoint; reduce size to find boundary.
# IPv6 headers are 40 bytes, so do not reuse the arithmetic blindly.What to expect
A reproducible size boundary, correct overhead arithmetic and evidence of the relevant ICMP or retransmission behavior.
Your turn
Why can MSS clamping fail to repair a large UDP application?
Show answer and reasoning
MSS is a TCP negotiation field. UDP must fit the path or use appropriate application/transport fragmentation behavior; fix MTU and required ICMP handling.Watch for: A jumbo-frame setting at one end is not end-to-end jumbo support.
Lesson references
Link to this lesson20. Verify address delivery end to end
For DHCPv4 trace Discover, Offer, Request and Acknowledgment through the correct VLAN and relay. Record relay source/giaddr, scope selection, exclusions, lease capacity, reservation, options and server reachability. Broadcasts normally do not cross routers without a relay. A relay forwards requests; it does not guarantee the server selects the intended scope.
Check delivered mask, gateway, DNS servers, search suffix and any application-specific options. A valid-looking address from the wrong DHCP server can be worse than no address. DHCP snooping must trust the legitimate server/relay path and remain untrusted on ordinary client ports; incorrect trust placement can drop valid offers.
Test a genuinely new lease and renewal where permitted, preserving the original evidence. Static hosts need separate address-conflict and option checks. APIPA/link-local IPv4 suggests failed automatic configuration but does not locate the failure. For DHCPv6, remember that Router Advertisements provide default-router information.
Windows:
ipconfig /all
Linux (available tools vary):
ip address
journalctl -u NetworkManager --since '10 minutes ago'
Capture filters: udp port 67 or udp port 68
Expected sequence: D → O → R → A, with transaction IDs correlated.What to expect
An approved server supplies the correct scope and options, with lease/relay/capture evidence.
Your turn
Discover reaches the relay, but no Offer returns. What boundaries do you inspect?
Show answer and reasoning
Relay-to-server routing/ACLs, server scope selection/capacity, server logs, return route and relay/client-side delivery. Correlate the transaction rather than restarting the service first.Watch for: Releasing an active production lease can interrupt service; use a test endpoint or an approved window.
Link to this lesson21. Separate authoritative, recursive and client DNS
A client normally asks a recursive resolver, which may use cached data or query authoritative servers. Test the resolver the application actually uses, not only a public resolver. Split DNS, VPN search domains, conditional forwarding, hosts files and browser/application DNS-over-HTTPS can produce different answers on one machine.
Record query name, type, resolver, response code, answer, TTL and time. NXDOMAIN means the name is reported absent; SERVFAIL means resolution failed; a timeout is no response. A/AAAA, CNAME, SRV and PTR records serve different purposes. A correct A record does not prove the AAAA path works, and a CNAME target can introduce another failure.
DNS uses UDP and TCP. Large responses, truncation, DNSSEC validation and broken TCP/53 paths can make only some names fail. Test fully qualified names before blaming a search suffix. After a change compare authoritative truth, recursive cache and client cache before clearing anything.
dig @192.0.2.53 app.example.net A
dig @192.0.2.53 app.example.net AAAA
dig @192.0.2.53 app.example.net A +tcp
Windows:
Resolve-DnsName app.example.net -Server 192.0.2.53
Resolve-DnsName _service._tcp.example.net -Type SRV -Server 192.0.2.53
# example.net is documentation; use your authorized lab zone.What to expect
Correct answers from intended resolvers, with both IPv4/IPv6 and TCP fallback considered.
Your turn
One resolver returns an old address while the authoritative server returns the new one. What do you inspect?
Show answer and reasoning
Remaining TTL, negative/positive caching, forwarding path, split-horizon view and whether the query really reaches the expected resolver.Watch for: Changing DNS servers can bypass an internal zone and create a new failure rather than solve the original one.
Link to this lesson22. Treat time, identity and certificates as dependencies
Authentication, logging correlation and certificate validation rely on accurate time. Inspect actual synchronization state, source, offset and reachability rather than merely finding an NTP server configured. Time zones affect display; UTC timestamps with offsets make cross-system evidence comparable.
For TLS inspect the requested hostname/SNI, certificate subject alternative names, chain, trust store, validity period and any interception proxy. A successful TCP/443 handshake does not establish TLS trust or application authorization. Test the intended hostname; connecting by IP can legitimately fail name validation or reach a different virtual host.
For enterprise identity verify DNS/SRV discovery, time, directory reachability, certificate enrollment and device/user policy. RADIUS/TACACS+ failures can be transport, shared-secret/certificate, identity, authorization or accounting problems. Preserve a controlled local recovery account according to the organization’s policy and test failover without locking out all administrators.
Linux:
timedatectl status
chronyc tracking
chronyc sources -v
openssl s_client -connect app.example.net:443 -servername app.example.net -verify_return_error
Windows:
w32tm /query /status
Test-NetConnection app.example.net -Port 443What to expect
Synchronized clocks, valid hostname-specific trust and a working identity path with documented recovery.
Your turn
The certificate is valid but the client reports it is not yet valid. What do you compare?
Show answer and reasoning
Client/server time and timezone display, the actual presented certificate, interception and chain validity. Do not disable verification as a permanent fix.Watch for: Certificate-private-key material belongs in secure storage, not a general evidence attachment.
Link to this lesson23. Validate wireless beyond association
Association with an SSID proves only part of the path. Inspect RF coverage, channel plan, channel width, interference, client capabilities, signal/noise and airtime utilization. Received signal strength alone does not measure available capacity. Excessively wide channels can increase interference and reduce channel reuse.
Follow authentication and addressing: SSID → security method → RADIUS/certificate if used → VLAN assignment → DHCP → DNS → gateway → application. Dynamic VLAN or role assignment can place a correctly authenticated user in the wrong segment. AP management and client forwarding can use different paths, especially with controller tunnels.
Roaming depends on client decisions, RF overlap, authentication behavior and supported fast-transition features. Test a real voice/video or application session while moving, not just whether the icon stays connected. Record AP/BSSID, band, channel, negotiated rate and actual throughput; PHY rate is not application throughput.
Wireless evidence:
Client MAC (randomized or stable?), SSID, BSSID, AP and switchport
Band/channel/width, RSSI/SNR, retries and utilization
Authentication result and assigned role/VLAN
Lease, DNS, gateway and application result
Roam timestamps and observed session interruptionWhat to expect
Approved clients authenticate into the correct segment and maintain required service quality at representative locations.
Your turn
Excellent signal but poor throughput affects all clients on one AP. What competing causes fit?
Show answer and reasoning
Airtime contention/interference, excessive retries, low-rate clients, uplink congestion, controller path, rate policy or endpoint/server limits. Compare RF and wired-side evidence.Watch for: A new SSID or open security setting is not a neutral troubleshooting change.
Link to this lesson24. Measure congestion and apply QoS deliberately
QoS chooses how constrained resources treat traffic; it does not create bandwidth. Identify the actual bottleneck and classify traffic at a trusted boundary. End-user DSCP markings may be untrusted. Define how markings map to queues and how the provider treats or rewrites them.
Shaping buffers traffic to a rate; policing drops or remarks traffic exceeding a profile. A priority queue can protect delay-sensitive traffic but must be bounded so other classes are not starved. Voice depends on latency, jitter and loss, while bulk transfers are often more tolerant of delay. Set targets from the application requirement.
Observe queue counters during representative load. A policy attached to the wrong direction/interface may show no useful hits. Test competing traffic classes together; an idle-network ping cannot demonstrate QoS effectiveness. Include provider CIR, tunnel overhead and encapsulation marking behavior in the design.
Cisco IOS-like:
show policy-map interface
show interfaces
# Record offered, transmitted, queued, dropped and remarked counters by class.
# Compare before/after deltas under the same controlled traffic mix.
# Check outer/inner DSCP behavior for tunnels on the exact platform.What to expect
Under agreed load, critical applications meet targets and the observed queue behavior matches policy.
Your turn
A priority policy is configured but voice still suffers at a 100 Mb/s provider policer on a 1 Gb/s interface. Why?
Show answer and reasoning
The local interface may never congest, so its queues do not control the provider bottleneck. A correctly sized parent shaper and class policy may be needed, accounting for the provider’s rate/overhead definition.Watch for: Copying a DSCP table without understanding trust and queue capacity is not a complete QoS design.
Lesson references
Link to this lesson25. Measure throughput, latency, loss and jitter honestly
Separate line rate, IP throughput and application goodput. Ethernet/IP/TCP/TLS overhead, endpoint CPU, storage, encryption, window size, RTT, loss and parallelism all affect results. The bandwidth-delay product estimates data needed in flight: 100 Mb/s × 0.080 s = 8 Mb, or 1 MB. A small receive window can limit a long-RTT transfer despite an uncongested circuit.
Record test endpoints, tool/version, direction, duration, stream count, payload size, rate cap, interface counters and CPU. Compare one stream with several only when the requirement justifies it; multiple streams can hide a single-flow limit. UDP tests report loss/jitter at the offered rate but can overload a link if unbounded.
Use isolated labs for overload-oriented benchmarking. RFC 6815 explicitly limits RFC 2544-style benchmarking to isolated environments; do not run an indiscriminate line-rate benchmark through shared production resources. For live-service verification use an agreed bounded method, relevant service criteria and provider coordination.
Authorized isolated lab, iperf3 installed on both endpoints:
Server: iperf3 -s
Client: iperf3 -c 198.51.100.20 -t 15 -P 1
Reverse: iperf3 -c 198.51.100.20 -t 15 -P 1 -R
Bounded UDP lab: iperf3 -c 198.51.100.20 -u -b 5M -t 15
# These generate traffic; choose a safe endpoint/rate and do not expose an open public test server.What to expect
A reproducible result with endpoint and network constraints identified, compared to an agreed target.
Your turn
Why might a speed test with eight streams pass while one file transfer remains slow?
Show answer and reasoning
Parallel streams can bypass a single-flow window, loss, hashing or server limitation. Test the actual workload and inspect RTT, retransmissions, window scaling and endpoint resources.Watch for: A test result without direction, load, duration and endpoint capability is not a trustworthy capacity claim.
Lesson references
Link to this lesson26. Use packet evidence without losing context
Choose capture points to answer a specific question: did the client transmit, did the firewall forward, did the server reply, and did the response return? Synchronize clocks and record interface/VRF, direction, snap length, filter and capture duration. SPAN can drop mirrored traffic under oversubscription, so absence in one capture is not automatically absence on the wire.
Interpret TCP sequence/acknowledgment progress, retransmissions, resets, window advertisements and TLS stage. A reset’s apparent source can be a middlebox; correlate TTL/MAC/path evidence rather than assuming the named host generated it. NIC checksum offload can make locally captured outgoing packets appear to have bad checksums.
Capture only the endpoints/ports needed and avoid collecting payload unnecessarily. Ring buffers and bounded duration prevent uncontrolled storage growth. Retain the original securely and share a redacted summary where possible. Do not upload sensitive packet captures to a public analysis service.
Authorized Linux capture:
sudo tcpdump -ni eth0 -s 128 -c 200 'host 198.51.100.20 and tcp port 443'
Wireshark display filters (not tcpdump capture syntax):
ip.addr == 198.51.100.20 && tcp.port == 443
tcp.flags.syn == 1
tcp.analysis.retransmission
icmp || icmpv6
# Capture filters and display filters are different languages.What to expect
Correlated ingress/egress observations answer the stated question with limitations recorded.
Your turn
A local capture flags bad TCP checksums on every outgoing packet but remote traffic works. What hypothesis comes first?
Show answer and reasoning
Checksum offload artifact. Compare a wire-side/remote capture and NIC offload behavior before diagnosing actual corruption.Watch for: Capture is an observation point, not omniscience; mirror loss and filters can hide packets.
Link to this lesson27. Trace virtual, container and cloud paths
A guest’s virtual NIC connects through a vSwitch/bridge, port group, host uplink and physical network. Verify VLAN tagging responsibility: guest-tagged traffic on an access-style port group can be double-tagged or dropped. Check hypervisor teaming, MTU, security controls and the physical switch’s LACP expectations; not every virtual teaming mode is LACP.
Containers add namespaces, bridges/overlays, NAT and policy. A service address can be translated to a changing backend, and a healthy pod does not prove Service selectors or network policy allow access. Trace client → service/load balancer → backend → return path. Keep the distinction between control-plane health and data-plane delivery.
Cloud routing includes route tables, security groups, network ACLs, gateways, peering/transit and managed DNS. Stateful and stateless policy semantics differ. Peering is not necessarily transitive. An established private circuit or VPN does not automatically advertise every cloud subnet or authorize every flow.
Read-only examples in the correct environment:
Linux: ip link; ip route; bridge vlan show
Kubernetes:
kubectl config current-context
kubectl get services,endpointslices -A
kubectl get networkpolicy -A
# Inspect selectors/readiness and CNI-specific policy with authorized access.
# Do not assume a cloud security group and network ACL have identical statefulness.What to expect
The full virtual-to-physical/cloud path and return path are documented, with each policy and translation boundary identified.
Your turn
A VM reaches peers on its host but not another host in the same VLAN. Where do you focus?
Show answer and reasoning
Host uplink/port-group VLAN carriage, physical trunk, teaming/LACP and MTU, while comparing guest addressing. The local vSwitch path already works.Watch for: Changing guest firewall policy cannot fix a VLAN missing from the hypervisor uplink.
Link to this lesson28. Handle multicast, discovery and specialized traffic
Broadcast discovery is usually limited to a subnet. Multicast uses group membership and, across subnets, multicast routing. IGMP/MLD snooping controls Layer 2 delivery but needs correct querier behavior; without it, memberships can age out or traffic can flood. PIM and reverse-path forwarding checks introduce a separate multicast control plane.
For a multicast application record source, group, receiver VLANs, source-specific versus any-source model, TTL/hop limit and expected rate. Verify membership, snooping state, querier, multicast route and RPF interface. A working unicast ping does not establish multicast forwarding.
mDNS/Bonjour, SSDP, PXE, voice and industrial/discovery protocols may depend on local broadcasts, relays or specific gateways. Extend discovery only where the design requires it; indiscriminate reflection across security zones can expose devices and create load. PXE also depends on DHCP options or proxy services and architecture-specific boot files, not merely TFTP reachability.
Cisco-like, platform-dependent:
show ip igmp groups
show ip igmp snooping groups
show ip pim neighbor
show ip mroute
show ip rpf 192.0.2.10
# Compare expected source/group and incoming/outgoing interfaces.What to expect
Only intended receivers obtain the stream, with stable membership and correct routed forwarding.
Your turn
A multicast stream works briefly after a switch restart then disappears. What state might be aging?
Show answer and reasoning
IGMP/MLD snooping membership due to missing or unreachable querier, alongside application join behavior. Inspect timers and queries rather than disabling snooping blindly.Watch for: A discovery relay is a deliberate security and scaling decision, not a universal connectivity fix.
Link to this lesson29. Make faults observable before they happen
Collect interface state/errors/discards, utilization, CPU/memory, temperature/power, routing neighbors, VPN status, DHCP/DNS health and application transactions. Distinguish polling, traps/events, streaming telemetry, flow records and logs. Each offers different timing and detail. SNMPv3 and authenticated/encrypted management should follow the platform and organizational policy.
Set thresholds against baseline and service impact. A high utilization alarm can be normal for a scheduled transfer; a low average can hide short microbursts and queue drops. Include loss of telemetry as an alert. Ensure monitoring reaches devices through intended paths and does not depend entirely on the same failing service.
Synchronize clocks, define retention and test a controlled fault. Confirm the alert arrives with device/interface/service context and a useful runbook. A dashboard showing green after collectors stopped is worse than an explicit unknown state.
Minimum baseline record:
Device/serial/software, interface role, speed, MTU
5-minute and short-interval utilization where available
Error/discard deltas, queue drops, link transitions
Neighbor/session uptime, CPU/memory, temperature/power
Synthetic DNS + TCP/TLS + authenticated application result
Collector last-seen timestamp and alert routeWhat to expect
A controlled failure produces an actionable alert and recovery signal; stale data is visibly stale.
Your turn
Average utilization is 20% but voice experiences bursts of loss. What do you add?
Show answer and reasoning
Shorter-interval or queue-level telemetry, microburst indicators, class drops, packet timing and correlated application metrics. Averages can hide transient congestion.Watch for: Monitoring credentials and flow records can expose sensitive topology; scope access and retention.
Link to this lesson30. Use repeatable configuration and evidence collection
Keep intended configuration, templates and inventories versioned. Separate secrets from ordinary source. Render device-specific candidates, validate syntax/schema and review diffs before applying. Idempotence means rerunning converges on the intended state; it does not mean an incorrect intended state is safe.
For APIs record endpoint/version, authentication scope, timeouts, pagination, rate limits and error handling. A 200 response can still contain an application-level error or partial result. For configuration deployment, use a small canary, verify service outcomes and stop on unexpected differences. Back up the prior state and understand rollback semantics.
Automate read-only evidence first: inventory, interface counters, neighbors, routes and timestamped outputs. Normalize device names and time zones so comparisons are meaningful. Do not blindly apply a generated command to every platform; unsupported syntax and differing defaults are common failure sources.
Change pipeline:
Inventory → render candidate → schema/syntax check → diff review
→ one lab/canary → functional verification → bounded rollout
→ post-checks → archive evidence
Evidence key:
UTC timestamp + device + command/API version + exit/status + output hash
Secrets: vault/reference, never plaintext in the reportWhat to expect
Repeatable candidates and evidence with explicit failure handling and a tested restoration path.
Your turn
A script exits zero but only half the switches changed. What should the workflow have checked?
Show answer and reasoning
Per-device API/command results, timeouts, partial failures, post-change state and functional outcomes. Process exit alone is not fleet success.Watch for: Automatic retries can repeat a non-idempotent mutation; know the operation’s semantics.
Link to this lesson31. Use a symptom-to-evidence fault matrix
Start with scope: one application, one host, one VLAN, one location, one address family or everyone? Record the last known good time and recent changes without assuming they caused the fault. Reproduce from a known source and compare a working peer. Change one variable at a time.
For no link, inspect physical/admin state and power. For no lease, trace DHCP. For gateway failure, inspect mask/VLAN/ARP/ND and gateway state. For remote-IP failure, inspect routes/policy/return path. For name-only failure, inspect resolver selection and answers. For TCP success but application failure, inspect TLS, proxy, identity and application logs. For large-only failure, inspect MTU. For intermittent failure, correlate counters, rekeys, leases, failovers and load.
Write hypotheses that can be falsified. “Firewall issue” is too vague; “the return packet is dropped at FW-A because the session entered FW-B” predicts a session/capture difference. Preserve evidence before restarting. A restart may restore service while erasing the cause; distinguish restoration from root-cause proof.
Observation → candidate causes → discriminating test
Small works, large stalls → MTU / endpoint / inspection → size sweep + captures
One VLAN fails → tagging / gateway / policy → MAC path + SVI + rule counters
IPv4 works, IPv6 fails → RA / IPv6 route/policy/DNS → explicit v6 transaction
Periodic outage → rekey / lease / scheduled load → timestamp correlation
Established TCP, TLS error → trust/name/time/interception → certificate + logsWhat to expect
A narrowed fault domain and evidence-backed explanation, with remaining uncertainty stated.
Your turn
Traceroute shows 80% loss at hop 4 but the destination has 0% loss. Is hop 4 dropping transit traffic?
Show answer and reasoning
Not demonstrated. The router may rate-limit responses to probes while forwarding transit traffic normally. End-to-end evidence and correlated captures matter.Watch for: Correlation with a recent change is a useful lead, not proof of causation.
Link to this lesson32. Test failures, recovery and hidden dependencies
List failure scenarios before pulling a cable: one link, one member, one switch, one power feed, one gateway, one carrier, one VPN peer, one resolver and one identity server. Predict the expected path and service impact for each. Verify independence first; two devices may share power, fiber route, configuration error or a control-plane dependency.
Measure detection time, convergence, lost packets/transactions, existing-session behavior and restoration. Failback can be more disruptive than failover, especially with preemption or asymmetric stateful paths. Test a representative authenticated application and long-lived session, not just continuous ping.
Run disruptive tests only in an isolated lab or an agreed window with rollback and observers. Restore the initial state, confirm redundancy is healthy again and retain the evidence. A backup path left active with the primary still broken is a restored service, not a completed resilience test.
Failure record:
Scenario and dependency removed
Predicted alternate path and acceptance threshold
Start UTC, detection UTC, recovery UTC
Packet loss + application transaction/session result
Failback result and final primary/backup state
Unexpected dependency and corrective actionWhat to expect
Measured behavior for each required failure, including failback and restoration of redundancy.
Your turn
Two routers fail together despite separate racks. What shared dependencies do you inspect?
Show answer and reasoning
Power upstream, carrier/physical route, shared switch/firewall, control plane, software/configuration, DNS/identity and monitoring path.Watch for: Redundancy counts components; resilience requires observing the service under failures.
Link to this lesson33. Trace 802.1X, NAC and endpoint admission
802.1X involves a supplicant (endpoint), authenticator (switch/AP) and authentication server, commonly reached through RADIUS. The endpoint exchanges EAP over the local link; the network device relays the exchange to the server. A link can be physically up while the controlled port remains unauthorized. Record the session state before treating the problem as a routing fault.
For EAP-TLS verify client and server certificates, trust chains, expected identity, time, revocation reachability where required and supported EAP configuration. A successful authentication can still produce a restricted VLAN, downloadable ACL or role. Inspect the authorization attributes and compare the client’s resulting VLAN/address with policy. Machine and user authentication can have different results at boot, login and reauthentication.
MAC Authentication Bypass is a separate fallback with weaker identity assurance; do not enable it broadly just to make a test pass. Phones with attached PCs, printers and other nonsupplicant devices need explicit multi-domain/multi-auth behavior. Decide what happens when the RADIUS server is unreachable, and test that policy without silently turning every failure into unrestricted access.
Evidence chain:
physical link → EAPOL exchange → RADIUS request/response
→ authenticated identity → assigned role/VLAN/ACL → DHCP → application
Cisco-like read-only example, platform-dependent:
show authentication sessions interface Gi1/0/10 details
# Newer platforms may use show access-session instead.What to expect
The correct identity receives the intended access, and rejection/server-unreachable cases follow documented policy.
Your turn
RADIUS says Access-Accept but the user cannot reach the app. What next?
Show answer and reasoning
Authorization attributes, dynamic VLAN/ACL, switch session state, DHCP in the assigned VLAN, routes and application policy. Authentication success is not end-to-end authorization proof.Watch for: Bypassing certificate validation or opening the port permanently changes the security design.
Link to this lesson34. Separate management, control and forwarding planes
The management plane handles administrative sessions and telemetry; the control plane computes routes and protocol state; the data plane forwards traffic. A device can forward packets while its CPU or management service is unhealthy, and can answer SSH while hardware forwarding is broken. Diagnose the plane associated with the symptom.
Restrict management sources and use a dedicated management VRF or out-of-band network where required. Use SSH/HTTPS, role-based access, AAA, audit logs and platform-supported secure telemetry. Define a tested emergency local-access procedure and protect its credentials. Disable unused management services and verify both IPv4 and IPv6 exposure.
Control-plane policing protects routing protocols and device-address traffic but can also drop legitimate probes or routing sessions if misconfigured. High CPU requires process and interrupt evidence, not only a percentage. Look for storms, route churn, excessive logs, scans, telemetry load or hardware punt reasons. Keep infrastructure addresses and protocol neighbors documented so legitimate control traffic can be distinguished from unexpected traffic.
Read-only evidence, exact commands vary:
show processes cpu sorted
show processes memory
show logging
show users
show control-plane
# Inspect platform-specific CPU queues/punt reasons and CoPP counters.
# Test access from one approved and one unapproved source in the lab.What to expect
Administrative access is controlled and recoverable; plane-specific evidence explains device health.
Your turn
Pings to a router drop but traffic through it is healthy. What distinction matters?
Show answer and reasoning
Packets destined to the router may be CPU-processed or rate-limited, while transit traffic is hardware-forwarded. Compare control-plane policy and end-to-end service before declaring a forwarding outage.Watch for: An administrative ACL applied to the wrong direction can lock out the recovery path.
Link to this lesson35. Separate WAN underlay from SD-WAN policy
The underlay provides basic transport reachability through a carrier, broadband or cellular connection. The overlay creates tunnels and a policy-controlled topology over it. Controller registration, tunnel state, route distribution and actual data forwarding are separate checkpoints. A controller dashboard may show the device connected while one application path is unusable.
Record each transport’s addressing, NAT/CGNAT, MTU, bandwidth, loss/latency profile, public/private reachability and provider constraints. Check controller/PKI time and certificate dependencies. For cellular backup include signal quality, antenna placement, data limits, address changes and whether inbound initiation is possible. Never assume a second transport is independent if both share the same building entry or upstream carrier.
Application steering can select paths by SLA probes, policy, business intent or learned application class. Probe destinations and intervals must represent the intended service; a healthy probe to a nearby gateway does not prove a remote SaaS path. Test brownouts (loss/jitter), not only complete link-down, and record existing-session behavior when a path changes. Distinguish policy failover from routing convergence and tunnel recovery.
Per-transport record:
carrier/service ID, interface, IP/NAT, gateway, MTU, contracted rate
controller/control sessions, overlay tunnel IDs, route membership
SLA probe destination and observed latency/loss/jitter
chosen app path, policy reason, failover/failback transaction resultWhat to expect
The chosen path and fallback behavior match the policy under normal load, degradation and restoration.
Your turn
Both tunnels are up but traffic never leaves the preferred path despite severe loss. What do you compare?
Show answer and reasoning
SLA probe coverage/thresholds, application classification, policy order, alternate-path eligibility and whether policy applies to existing sessions or only new flows.Watch for: Adding a cellular link does not automatically provide inbound reachability or carrier diversity.
Link to this lesson36. Check proxies, load balancers and application dependencies
A reverse proxy or load balancer can terminate TCP/TLS and create a separate connection to a backend. Treat the client-to-proxy and proxy-to-server paths as distinct flows with separate DNS, certificates, routes, policies and source addresses. Health checks may use a different path, Host header or authentication method from real users.
For virtual-hosted HTTPS retain the intended hostname and SNI. A bare-IP request can reach the default virtual host and generate a misleading certificate or 404 error. For explicit proxies inspect client proxy/PAC selection, proxy authentication, CONNECT policy and trust of any inspection certificate. Transparent inspection creates different evidence again.
A green TCP health check does not prove database, identity or storage dependencies. Use a meaningful application transaction with a nonprivileged test identity and known expected response. For long-lived sessions consider persistence, idle timeout, backend draining and failover. For directory/file services, distinguish name discovery, transport, authentication and application authorization instead of opening broad port ranges without an application requirement.
Controlled hostname test, replace with your authorized lab endpoint:
curl -v --connect-timeout 5 --max-time 15 https://app.example.net/health
# Preserve hostname while targeting a known lab IP:
curl --resolve app.example.net:443:198.51.100.20 https://app.example.net/health
# Do not use -k as the permanent acceptance test.
# Keep tokens/cookies out of shared verbose output.What to expect
The real client transaction succeeds through the intended proxy/backend with valid identity and trust.
Your turn
The backend health check passes but authenticated requests fail. What can differ?
Show answer and reasoning
Host/SNI, URL, credentials, headers, persistence, TLS trust, backend dependency and path/policy. Reproduce the actual request safely rather than equating a TCP connect with application health.Watch for: Changing a health check to always return success hides a fault.
Link to this lesson37. Escalate with a bounded, reproducible evidence pack
An effective escalation names the service impact, fault scope, start time, known working comparison, suspected boundary and requested action. Include circuit/device identifiers, topology fragment, exact source/destination/VRF, relevant counter deltas and timestamped tests. State what has and has not been changed. Keep fact, hypothesis and assumption visibly separate.
For a carrier, include demarc status, optical/physical evidence, encapsulation, next-hop reachability and a test to an agreed endpoint. For a firewall team, include tuple, zone, NAT stage, session/rule evidence and return-path observations. For an application team, include successful network/TLS stages and the application error or transaction ID. Give each team evidence at its boundary rather than assigning blame from a generic timeout.
Maintain one incident timeline and change log. Record temporary workarounds, expiry/rollback owner and the effect on monitoring/security. During recovery compare the same test from the original failing source. Close root-cause work only when the explanation fits the observations; restored service alone may leave the cause uncertain.
Escalation template:
Impact and scope:
UTC onset / last known good:
Exact failing flow and working comparison:
Boundary evidence and attachment references:
Recent changes / actions already attempted:
Observed facts:
Hypothesis and discriminating test requested:
Owner, next update and workaround expiry:What to expect
Another team can reproduce or investigate the precise boundary without first reconstructing basic facts.
Your turn
How do you report a restart that restored service but erased logs?
Show answer and reasoning
Service restored after restart; root cause unconfirmed. Record the pre-restart observations, recovery time and missing evidence, then define monitoring for recurrence.Watch for: Do not send full configurations or captures when a redacted, scoped extract answers the question.
Link to this lesson38. Worked incident: a small ping passes but the application stalls
Observation: a client reaches its gateway and receives small ping replies from the application server. TCP/443 establishes, but a larger authenticated response repeatedly stalls after a WAN tunnel change. This supports several hypotheses: reduced path MTU with blocked feedback, application/backend delay, TLS inspection or endpoint limits. The successful ping alone does not choose among them.
First preserve the exact client address, resolver answer, routing domain, time and application error. Confirm the same server/hostname works from a comparison path. Capture a bounded failing transaction at the client and tunnel boundary. Repeated transmission of the same larger TCP segment with no progress, combined with a reproducible DF-packet size threshold and missing ICMP feedback, supports an MTU black-hole diagnosis. A server that never sends the response would instead shift attention toward application dependencies.
Calculate the tunnel’s actual overhead for the selected encapsulation and compare inner/outer MTU. Restore required ICMP handling and the designed tunnel MTU/MSS policy through the approved change, then retest small and large transactions, both directions, IPv4 and IPv6 if in scope. Do not declare victory from a new ping only. Record the packet-size boundary and application result before/after, plus the permanent monitoring or configuration check that will catch recurrence.
Evidence sequence:
1. Same hostname/IP/source/VRF recorded
2. TCP handshake succeeds
3. Large response retransmits without ACK progress
4. Smaller controlled packet succeeds; larger DF packet fails
5. Tunnel overhead + missing feedback explain boundary
6. Approved correction
7. Original authenticated transfer succeeds; negative policy tests remain intactWhat to expect
A diagnosis supported by multiple independent observations and verified through the original workload.
Your turn
What observation would weaken the MTU hypothesis most strongly?
Show answer and reasoning
The server never emits the application response despite receiving the complete request, while controlled large transfers across the same path succeed. Investigate backend/application dependencies instead.Watch for: The scenario is a worked reasoning example, not a claim that every “ping works, app fails” incident is MTU.
Link to this lesson39. Close with evidence and a usable handover
Acceptance is a traceable comparison with the requirements established at the start. Each result needs source/destination, protocol, time, expected outcome, measured outcome and evidence. Mark pass, fail, not tested or accepted exception explicitly. Do not turn “not tested” into pass because the deadline arrived.
Deliver as-built physical/L2/L3/security diagrams, port schedule, addressing/IPAM, circuit identifiers, device/software inventory, configuration backup locations, monitoring/runbooks, support contacts, license/support renewal information and known issues. Record secrets by secure reference rather than embedding them. Include exact recovery steps and tested access methods.
Observe the service for an agreed stabilization period with representative traffic. Transfer ownership with a named receiving operator and unresolved actions. An incident-ready handover should let another engineer locate a port, identify a circuit, understand a route/policy, collect the right evidence and restore service without reconstructing the entire design.
Acceptance row:
ID | requirement | source/destination | test | target | measured
| UTC | evidence | status | owner | exception expiry
Handover essentials:
as-built + inventory + IPAM + circuits + secure backups
monitoring + recovery runbooks + support/escalation + known issuesWhat to expect
A receiving engineer can independently trace, verify and recover the service using the handover.
Your turn
Failover was never tested, but all normal-path tests pass. How should acceptance represent it?
Show answer and reasoning
Normal-path tests pass; failover remains not tested or an explicitly accepted exception with owner and due date. Do not imply resilience was demonstrated.Watch for: A polished report cannot substitute for missing test evidence.
Link to this lesson40. Integrated lab: build, break, diagnose and restore
Use an isolated two-router/two-switch lab with two client VLANs and a separate management VLAN. Use the documentation ranges in the address plan, a routed transit link and an application server on 198.51.100.0/24. Keep the lab disconnected from real enterprise/carrier networks. Packet Tracer, virtual appliances or physical hardware can be used, but unsupported features must be marked rather than simulated as passing.
First configure and verify physical links, VLANs/trunks, gateways and static routing. Add DHCP relay/server and DNS for a lab application. Apply a policy permitting the approved client-to-server flow and denying guest-to-management. Capture a baseline transaction. Then add a redundant path using the platform’s supported design and measure failure/recovery.
Inject one fault at a time: remove a VLAN from a trunk; use a wrong client mask; remove the return route; change a DNS answer; block required PMTU feedback; introduce a wrong VPN selector if your lab supports VPN. Before opening the answer, write the predicted symptom and the observation that would distinguish it from another cause. Restore the baseline after each fault and repeat the same application test.
Finish with an evidence pack. A successful lab includes negative tests and a usable rollback, not just screenshots of configuration. Do not claim physical optic, RF, provider policer or hardware failover validation from a simulator that does not model those behaviors.
Lab address fixture:
Users: 192.0.2.0/26, gateway .1, test client .10
Guests: 192.0.2.64/27, gateway .65, test client .70
Management: 192.0.2.96/28, gateway .97
Transit: 192.0.2.112/30, routers .113 and .114
Server LAN: 198.51.100.0/24, gateway .1, app .20
DNS: choose a lab server and document it; no public records needed
Baseline tests:
client → own gateway; client → app by IP; client → app by name
approved TCP transaction; forbidden guest → management attempt
reverse route; counter deltas; controlled failure; restoreWhat to expect
A complete isolated topology with repeatable positive/negative tests, fault explanations and a restored baseline.
Your turn
The app works by IP but not by name; a new wrong DNS answer was injected. What evidence proves the intended fault rather than merely suggesting it?
Show answer and reasoning
Query the actual client resolver, record the wrong answer/TTL, compare the authoritative intended record, show successful IP transport and hostname/TLS consequences, restore the record/cache state deliberately and repeat the transaction.Watch for: A lab is a learning environment; copy concepts and validated platform syntax, never its documentation IPs into a real service.
Link to this lessonReferences
Original AUWEN lessons, with upstream documentation for further study and version checks.
- IETF RFC index and errata ↗
- Wireshark User’s Guide ↗
- ESnet iperf3 documentation ↗
- Cisco configuration and support documentation ↗
- Juniper documentation ↗