Every time you make a video call, stream a film, or send a message, your data travels across networks that were originally built for very different purposes. The telephone network was designed for voice. The internet was designed for data. For decades these systems spoke different “languages,” and getting them to work together was a real engineering challenge. Protocol convergence is the answer to that challenge. It is the process of unifying separate communication methods so that voice, video, and data can flow smoothly over a shared infrastructure. To understand how this convergence works, we need to look at the switching techniques that move information, the signaling protocols that set up conversations, and the transmission standards that carry everything across the optical backbone.
Table of Contents
Different switching techniques
At the heart of any communication network is switching: the method by which data is moved from a sender to a receiver. Over the years, three main approaches have shaped how networks operate, and each one solves the problem differently.
Circuit switching
Circuit switching establishes a dedicated physical path between two endpoints before any communication begins, and this path stays reserved for the entire duration of the conversation. The traditional telephone network, including the Public Switched Telephone Network (PSTN) and ISDN, works this way. The advantage is consistent quality: once the circuit is set up, you get guaranteed bandwidth and a stable connection. The disadvantage is inefficiency. The path is held open even during silences in a phone call, which means capacity is wasted whenever no data is actually flowing. Circuit-switched systems give reliable service but waste capacity when traffic is mixed or bursty.
Packet switching
Packet switching takes a completely different route. Instead of reserving a path, it breaks data into smaller units called packets, each carrying addressing information, and sends them independently across the network. The internet uses this model with protocols like TCP/IP and UDP. As packets travel through switches and routers, they are received, buffered, queued, and forwarded. This makes far better use of available bandwidth because many users can share the same links. The trade-off is variable latency. Because packets are forwarded as resources become available, delays can fluctuate depending on network load, which is not ideal for time-sensitive traffic like a live voice call.
Cell switching
Cell switching emerged as a middle path that tries to capture the strengths of both earlier methods. The best-known example is Asynchronous Transfer Mode (ATM). Instead of variable-length packets, cell switching uses small, fixed-length cells of 53 bytes, where 5 bytes form the header and 48 bytes carry the payload. ATM is connection-oriented, meaning a path is requested and resources are reserved before transmission, much like circuit switching, but it still transmits data in cell form, much like packet switching. This is why ATM is often described as capturing the advantages of both circuit and packet multiplexing.
The rise of cell switching
To appreciate why cell switching became so important to convergence, it helps to understand the problem it was built to solve. Networks need to carry two broad categories of traffic. Real-time traffic, such as voice and video, must arrive within a strict time window, otherwise the call breaks up or the video freezes. Non-real-time traffic, such as file transfers and email, can tolerate delay but is sensitive to data loss. Earlier networks struggled to serve both well on a single infrastructure.
The fixed size of ATM cells is the key innovation here. Because every cell is exactly the same length, switching hardware always knows in advance how long each cell is and exactly when the last portion will arrive. This makes fixed-length switching far faster and simpler to design than variable-length packet switching, where each packet must be examined for its length. The predictability of fixed cells also reduces jitter, the variation in delay between data units, which is exactly what real-time multimedia needs.
This design directly enabled the concept of Quality of Service (QoS) as we understand it today. ATM was built to efficiently serve both delay-sensitive real-time services and loss-sensitive non-real-time data on the same network. It does this through service categories that prioritise traffic according to its needs. Constant Bit Rate (CBR) provides fixed bandwidth with low delay for voice and video, Variable Bit Rate (VBR) handles bursty traffic with defined throughput requirements, Available Bit Rate (ABR) offers a guaranteed minimum cell rate for data, and Unspecified Bit Rate (UBR) provides a best-effort service with no guarantees. By giving each traffic type its own treatment, ATM allowed a single network to carry a telephone conversation and a file download simultaneously, each receiving appropriate handling.
ATM emerged in the 1980s and 1990s as part of the Broadband Integrated Services Digital Network (B-ISDN) project, with the International Telecommunication Union writing the first specifications. While ATM was never fully realised as the single universal technology its designers imagined, and is less common in today’s internet core, its ideas about fixed-length switching and traffic prioritisation remain foundational to how converged networks manage mixed traffic.
Signaling protocols
Moving data is only half the story. Before any voice or video session can begin, the network needs to set up the call: locate the other party, agree on which media formats to use, ring the destination, and tear the session down when it ends. This work is handled by signaling protocols. In the world of Voice over IP (VoIP), three names dominate the discussion: SIP, H.323, and the closely related RTCP.
H.323
H.323 is a protocol suite developed by the ITU Telecommunication Standardization Sector (ITU-T). It was originally designed as a standard for real-time videoconferencing on local area networks and grew into a comprehensive framework for multimedia communication. H.323 is not a single protocol but a family that includes H.225.0 for call signaling, H.245 for exchanging capability information and managing media channels, and a Registration, Admission and Status (RAS) function that gives it a hierarchical structure. This structure is useful for large enterprises with multiple branches that need centralised, interconnected telephony. The trade-off is complexity. H.323 is generally considered the more complex of the major signaling standards, and it encodes messages in a compact binary format suited to both narrowband and broadband connections.
SIP
The Session Initiation Protocol (SIP) was created by the Internet Engineering Task Force (IETF) and is far simpler in design. Where H.323 grew out of the telephony tradition, SIP grew out of the internet tradition, using a text-based, request-and-response style similar to web protocols. A basic SIP exchange can be built from just a few request types such as INVITE, ACK, and BYE, alongside headers like To, From, Call-ID, and CSeq. SIP works together with the Session Description Protocol (SDP) to describe the media in a session. Although H.323 appeared first, SIP is now the dominant choice because it adapts more naturally to the internet environment and integrates easily with applications like presence and instant messaging.
RTCP and the role of media transport
Both SIP and H.323 are signaling protocols; they set up and manage sessions but do not carry the actual voice or video. That job belongs to the Real-time Transport Protocol (RTP). Working alongside it is the RTP Control Protocol (RTCP). Importantly, both SIP and H.323 rely on the same RTP/RTCP pair for real-time data transmission. RTCP does not transport media itself; instead it carries statistics and control information about the media stream, such as packet loss, delay, and jitter, allowing endpoints to monitor quality and adapt. This separation, where signaling, media transport, and quality monitoring each have their own protocol, is a clear example of convergence in action: independently designed components cooperating to deliver a single seamless call.
Transmission standards
All of this switching and signaling ultimately depends on a physical backbone that can carry enormous volumes of data over long distances. This is the domain of optical fibre and two closely related standards: SONET and SDH. They are the high-capacity transport containers that move both voice and data across carrier networks.
SDH and SONET explained
SONET (Synchronous Optical Network) is the standard developed in North America by the ANSI committee for synchronous data transmission over optical fibre. SDH (Synchronous Digital Hierarchy) is the international equivalent, developed by the International Telecommunication Union and used across most of Europe, Asia, and the rest of the world, including in India. Both were created to solve the same problem. Before them, combining circuits from different sources was difficult because each circuit ran at a slightly different rate and phase. SONET and SDH introduced synchronised clocking so that many different circuits could be transported within a single framing structure.
Comparing SONET and SDH
The two standards are best understood as variations on a common theme rather than rival technologies. With a few exceptions, SDH can be regarded as a superset of SONET. Their most visible difference is the basic unit of transmission. SONET begins with the OC-1 rate of 51.84 Mbps and uses Optical Carrier designations such as OC-3 and OC-12, while SDH begins with STM-1 at 155.52 Mbps and uses Synchronous Transport Module designations. The internal terminology also differs: what SONET calls a Virtual Tributary, SDH calls a Virtual Container. Despite these naming differences, the two share core functions including framing, error checking, link management, and synchronised operation.
What makes these standards so relevant to convergence is their neutrality. SONET and SDH are not communication protocols in themselves; they are general-purpose transport containers capable of carrying many kinds of traffic. This is precisely why they became the chosen vehicle for transporting ATM cells and later for Packet over SONET/SDH networking. Their bandwidth-flexible virtual containers can hold voice, data, or video equally well. They also offer self-healing ring architectures that can reroute traffic in milliseconds if a fibre is cut, a level of reliability essential for backbone infrastructure. In this way, the optical transport layer ties the whole convergence story together, carrying the switched cells and the signalled sessions across the network on a single unified foundation.
How it all comes together
Protocol convergence is not a single invention but the cumulative result of these layers working in harmony. Cell switching gave networks a way to handle real-time and non-real-time traffic with predictable performance and quality guarantees. Signaling protocols like SIP and H.323, supported by RTP and RTCP, gave them a way to set up, manage, and monitor multimedia sessions across formerly separate systems. SONET and SDH gave them a neutral, reliable optical backbone to carry it all. Together, these technologies dissolved the old boundaries between the voice network and the data network, allowing a single converged infrastructure to deliver the calls, streams, and messages we now take for granted.
What do you think? If ATM was technically capable of unifying voice, video, and data with quality guarantees, why do you think the simpler, best-effort model of packet-switched IP networks ultimately came to dominate the modern internet? And as networks continue evolving toward newer optical and software-defined technologies, which qualities of these earlier convergence standards do you think will remain most relevant?
References
- https://www.ninjaone.com/blog/circuit-switching-vs-packet-switching/
- https://www.cse.wustl.edu/~jain/cis788-95/ftp/atm_cong/index.html
- https://www.sciencedirect.com/topics/physics-and-astronomy/packet-switching
- https://www.broadbandsearch.net/definitions/asynchronous-transfer-mode
- https://networkencyclopedia.com/cell-in-atm/
- https://www.sciencedirect.com/topics/computer-science/asynchronous-transfer-mode
- https://discoveryengineering.net/blog/asynchronous-transfer-mode-atm/
- https://info.teledynamics.com/blog/voip-protocols-h.323-and-mgcp-as-alternatives-to-sip
- https://ieeexplore.ieee.org/document/4561304/
- https://techdifferences.com/difference-between-h-323-and-sip.html
- https://homel.vsb.cz/~voz29/files/voz_29.pdf
- https://www.techtarget.com/searchnetworking/definition/Synchronous-Optical-Network
- https://en.wikipedia.org/wiki/Synchronous_optical_networking
- https://lightyear.ai/tips/sonet-versus-sdh
- https://www.cisco.com/c/en/us/support/docs/optical/synchronous-optical-network-sonet/16180-sonet-sdh.html

Leave a Reply