How to Detect STP Route Oscillation with Transfer Controlled Messages

In legacy PSTN cores, signalling depends on quiet MTP3 decisions. When a link set or Signal Transfer Point cluster flaps between reachable and unreachable, the symptoms rarely show as a clean alarm. They look like jitter, sluggish call setup, a mobile user in regional Australia complaining the call drops on the second ring. Catching the pattern early is the difference between a quick fix and a week of escalation.

Telstra, Optus, and SIP trunks on the NBN share the same MTP3 routing tables. In Sydney or Perth, congestion is urban and predictable. Out past Dubbo or along the Stuart Highway, the same load becomes route oscillation because the bearer keeps changing state. Engineers need a fast way to confirm the suspicion before opening a vendor ticket.

Transfer Controlled messages are the standard mechanism an STP uses to tell adjacent points to slow down or stop sending toward a congested destination. When a destination oscillates, TFC traffic spikes in a recognisable rhythm, and a careful reader of the MTP3 headers can extract the destination, point code, and priority. That trail is what a good detection routine follows.

This piece walks through what oscillation looks like at the protocol level, how TFC messages expose it, and what a practical pipeline looks like. For a refresher, the message transfer fundamentals article covers the basics.

What STP Route Oscillation Looks Like in Practice

Route oscillation is a pattern, not a single event. An STP that flips a destination between available and unavailable states generates bursts of Transfer Prohibited, Transfer Allowed, and Transfer Restricted messages. Each flip produces a wave of TFC traffic toward signalling points forwarding user part messages through the questionable path.

In Australian backhaul, the trigger is often a flaky transmission link between a regional aggregation router and a core site. A team in Adelaide might see a stable route suddenly start flapping every 90 seconds, each flap well under a second. A useful first indicator is a sudden rise in TFC messages on a historically quiet destination. The asymmetry between an affected path and a healthy backup is the fingerprint.

The Role of Transfer Controlled in Flow Control

When an STP detects that a destination's congestion level has reached a defined threshold, it issues a Transfer Controlled message back toward the originator. The receiving point then stops, slows, or re-routes user part messages for the affected destination and priority. This protects downstream links when a transit switch loses half of its link set.

For detection, TFC messages are a gift. They carry the affected destination point code, the congested destination when they differ, and the priority field. When these arrive in bursts aligned with the route flap pattern, the relationship becomes clear. Counting TFCs per destination per minute often confirms a fault the management layer has dismissed as noise. A sliding window with a rate threshold is the simplest practical method.

Decoding MTP3 Status Fields and TFC Variants

The Transfer Controlled message sits inside the standard SS7 header. The routing label destination point code identifies the signalling point being asked to hold traffic. The congested point code identifies where the bottleneck actually sits. The two match when the STP is the bottleneck itself, the most common case during oscillation.

Australian networks carry point codes in both 14-bit and 24-bit formats. Telstra historically used 14-bit, while newer virtualised platforms lean toward 24-bit. A detection script has to handle both, or it will miss roughly a third of the signal in mixed environments. Priority bits also matter: a high volume of TFCs against ISUP is a stronger indicator of call-affecting oscillation than TFCs against SCCP alone.

Building a Capture Workflow for Australian Carriers

The first step is to tap a signalling link at a representative aggregation point. For Telstra and Optus engineers, this usually means a SPAN port in front of a core STP. For smaller providers running their own SIP-to-SS7 gateways, the capture point is often a Linux box using a raw socket or a DAG card.

Wireshark remains the workhorse for ad hoc analysis, and its MTP3 dissector decodes TFC messages cleanly. The relevant filter is mtp3.service_indicator == transfer-controlled. Saving the filtered stream to a pcap allows later parsing with tshark or a Python script using scapy. Detection should happen well before any retention-driven review, with a real-time parser pushing alerts to a paging system.

Timing Thresholds and Flap Detection Logic

Thresholds are the heart of any oscillation detector. The most common mistake is to set them based on theory rather than observed data. A reasonable starting point is to collect 14 days of TFC counts per destination during a known-good period, then compute the 95th percentile of the per-minute rate. Anything more than four times that baseline is worth a human look.

Flap detection adds a second layer. A destination is flapping if it transitions between available and unavailable states more than three times in 300 seconds. Time-of-day adjustments help in the Australian context. Peaks around business hours in Sydney and Melbourne push TFC counts higher, and a static threshold will generate false positives during large sporting events, when Triple Zero (000) load produces legitimate congestion that should not be misread as oscillation.

Separating Real Oscillation from Routine Churn

Not every rise in TFC volume is oscillation. Planned maintenance, a hardware swap in a remote exchange, or a wholesale shift in routing during an NBN cutover all produce bursts of TFC traffic that look similar at first glance. The detection routine needs to discount these cases.

The most reliable cross-check is the human ticket queue. If the network management system has an open change record for a destination matching the TFC spike, the alert can be suppressed automatically. Planned work almost always involves a deliberate link state change, while oscillation does not. The final sanity check is human, ensuring the engineer is looking at the right destination at the right time.

Comparing Detection Approaches

Approach Setup effort Best for Limitation
Manual Wireshark review Low One-off incidents Slow, error-prone at scale
tshark + cron script Medium Small teams, single STP No real-time alerting
Scapy-based real-time parser Medium-high Mixed 14-bit / 24-bit networks Requires Python skill on call
Commercial SS7 probe with TFC dashboards High Carriers, large ISPs Licensing cost, vendor lock-in

For most teams in Australia starting fresh, the second or third row is the practical choice. The fourth row makes sense once the operator is running more than a handful of STPs. Practitioners who want structured material on MTP3 decoding, timers, and call flow can find it in SS7 training programs that walk through the protocol layer by layer.

The lasting picture is straightforward: route oscillation leaves a fingerprint in TFC traffic, the fingerprint is consistent across 14-bit and 24-bit point code worlds, and a simple rolling-window script can catch it long before customers ring in. A good detector measures the right destination, the right priority, and the right window, then stays quiet long enough that the team trusts it when it does speak up. Build the threshold from real data, suppress the alerts that match planned work, and keep one engineer close enough to the script to recognise the patterns the algorithm cannot yet describe.