Queuing Theory and Operations Research

Queuing theory and operations research is a huge subject area and applies to so many situations. Queues are everywhere. At B-Agile, Our core interest is the identification and optimising of queues related to business processes like product, platform and service development, business operations and digital, AI transformation and change.
What is queuing theory?
Queuing theory refers to the mathematical study of the formation, function, and congestion of waiting lines, or queues. It’s also referred to as queueing theory, queue theory, and waiting line theory.
At its core, a queuing situation involves two parts:
- Someone or something that requests a service—usually referred to as the customer, job, or request.
- Someone or something that completes or delivers the services—usually referred to as the server.
To illustrate, let’s take two examples. First, when looking at the queuing situation at a bank, the customers are people seeking to deposit or withdraw money, and the servers are the bank tellers.
Second, when looking at the queuing situation of a printer, the customers are the requests that have been sent to the printer, and the server is the printer.
Queuing theory scrutinizes the entire system of waiting in line, including elements like the customer arrival rate, number of servers, number of customers, capacity of the waiting area, average service completion time, and queuing discipline.
Queuing discipline refers to the rules of the queue, for example whether it behaves based on a principle of first in first out, last-in-first-out, prioritized, or serve-in-random-order.
How did queuing theory start?
Queuing theory was first introduced in the early 20th century by Danish mathematician and engineer Agner Krarup Erlang.
Erlang worked for the Copenhagen Telephone Exchange and wanted to analyze and optimize its operations.
He sought to determine how many circuits were needed to provide an acceptable level of telephone service, for people not to be “on hold” (or in a telephone queue) for too long. He was also curious to find out how many telephone operators were needed to process a given volume of calls.
His mathematical analysis culminated in his 1920 paper “Telephone Waiting Times”, which contained some of the first queuing models and served as the foundation of applied queuing theory.
The international unit of telephone traffic is called the Erlang in his honour.

In queuing theory, all systems can be simplified into entities queuing for activities.
We can refer to this sub-system as managing customers queuing for service. To analyse it, we need information about:
Arrival process:
- How customers arrive (individually or in groups)
- The probability distribution of time between arrivals (interarrival time distribution)
- Whether the customer population is finite or infinite
Service mechanism:
- Resources needed to start the service
- Duration of the service (service time distribution)
- Number of servers available
- Server arrangement: series (separate queues) or parallel (one queue for all)
- Pre-emption policy: whether a server can interrupt a customer for an emergency
Queue characteristics:
- How do we choose the next customer to be served? Options include FIFO (first-in first-out), FCFS (first-come first-served), LIFO (last-in first-out), or randomly. This is known as queue discipline.
- Do we have:
- Balking (customers avoid joining a long queue)
- Reneging (customers leave if they wait too long)
- Jockeying (customers switch queues for faster service)
- Is the queue finite or infinite in capacity?
Queuing theory uses the Kendall notation to classify the different types of queuing systems, or nodes. Queuing nodes are classified using the notation A/S/c/K/N/D where:
- A is the arrival process
- S is the mathematical distribution of the service time
- c is the number of servers
- K is the capacity of the queue, omitted if unlimited
- N is the number of possible customers, omitted if unlimited
- D is the queuing discipline, assumed first-in-first-out if omitted
For example, think of an ATM. It can serve:
- One customer at a time
- In a first-in-first-out order
- With a randomly-distributed arrival process and service distribution time
- Unlimited queue capacity
- Unlimited number of possible customers.
Queuing theory would describe this system as a M/M/1 queuing model (“M” here stands for Markovian, a statistical process to describe randomness).
Most queuing systems in businesses are classified into one of these line forms

Why is queuing theory important?
Waiting in line is a part of everyday life because as a process it has several important functions. Queues are a fair and essential way of dealing with the flow of customers when there are limited resources. Negative outcomes arise if a queue process isn’t established to deal with overcapacity.
For example, when too many visitors navigate to a website, the website will slow and crash if it doesn’t have a way to change the speed at which it processes requests or a way to queue visitors.
Or, imagine planes waiting for a runway to land. When there is an excess of planes, the absence of a queue would have real safety implications as planes all tried to land at the same time.
The importance of queuing theory comes from the fact that it helps describe features of queues, like average wait time, and provides the tools for optimizing queues. From a business sense, queuing theory in operation research informs the construction of efficient and cost-effective workflow systems.
Queuing situations involve uncertainty in interarrival and service times, requiring probability and statistics for analysis.
In analysing queuing situations, we focus on system performance metrics like:
- Expected wait time before service and completion.
- Probability of waiting longer than a specific interval.
- Average queue length.
- Probability of the queue exceeding a certain length.
- Expected server utilisation and periods of full occupation. Assigning costs to customer wait time and server idle time can help design cost-effective systems.
Teams need to answer key questions to evaluate alternatives and improve situations. Common issues include:
- Is reducing service time worth the effort?
- How many servers are needed?
- Should customer priorities be set?
- Is the waiting area adequate?
To answer these questions, there are two main approaches: analytic methods (queuing theory) and simulation (computer-based).
Whilst queuing theory can be used to analyse simple queuing systems, more complex queuing systems are typically analysed using simulation (more accurately called discrete-event simulation).
Lets look at a simple example of a simulation:

Suppose that customers arrive with interarrival times that are uniformly distributed between 1 and 3 minutes, i.e. all arrival times between 1 and 3 minutes are equally likely. Suppose too that service times are uniformly distributed between 0.5 and 2 minutes, i.e. any service time between 0.5 and 2 minutes is equally likely.
We have two independent statistical distributions: interarrival times between 1 and 3 minutes, and service times between 0.5 and 2 minutes. We can generate lists of these times by sampling from the distributions using Excel formulas: `1+(3-1)*RAND()` for interarrival times and `0.5+(2-0.5)*RAND()` for service times.
Suppose our two lists are:
Interarrival times Service times
1.9 1.7
1.3 1.8
1.1 1.5
1.0 0.9
etc etc
Rounded to one decimal place
Consider our system at time zero (T=0) with no customers. What will happen next?
After 1.9 minutes, a customer arrives and is served immediately as the queue is empty and the server is idle.
Next, at T=3.2 (after 1.3 more minutes), another customer appears and joins the queue because the server is busy.
At T=3.6 (1.7 minutes later), the current customer finishes and leaves. The next customer in the queue starts their service, ending at T=5.4 (taking 1.8 minutes).
Another customer arrives at T=4.3 (1.1 minutes after the previous arrival) and joins the queue.
At T=5.3 (1.0 minutes later), another customer appears and also joins the queue, making it two customers waiting.
Finally, at T=5.4, the current customer finishes. The first customer in the queue begins their service, ending at T=6.9 (taking 1.5 minutes).
Obviously the above process is best done by a computer.
To summarise what we have done we can construct the list below:
Time T What happened
1.9 Customer appears, starts service scheduled to end at T=3.6
3.2 Customer appears, joins queue
3.6 Service ends
Customer at head of queue starts service, scheduled to end at T=5.4
4.3 Customer appears, joins queue
5.3 Customer appears, joins queue
5.4 Service ends
Customer at head of queue starts service, scheduled to end at T=6.9
We are employing discrete-event simulation to emulate our queuing system, focusing on specific events over time, such as customer arrivals and service completions at intervals like T=1.9, 3.2, 3.6, 4.3, 5.3, 5.4, etc.
From this simulation we can derive various statistics about the system, including the average duration a customer spends queueing and being served (average time in the system). In this example, two customers have completed the entire process: the first arrived at time 1.9 and departed at time 3.6, spending 1.7 minutes in the system. The second customer arrived at time 3.2 and departed at time 5.4, spending 2.2 minutes in the system. Consequently, the average time in the system is (1.7+2.2)/2 = 1.95 minutes.
Additionally, we can calculate statistics on queue lengths, such as the average queue size. Here, the queue size is 0 from T=0 to T=3.2, size 1 from T=3.2 to T=3.6, size 0 from T=3.6 to T=4.3, size 1 from T=4.3 to T=5.3, and size 2 from T=5.3 to T=5.4. Thus, the time-weighted average queue size is:
[0(3.2-0) + 1(3.6-3.2) + 0(4.3-3.6) + 1(5.3-4.3) + 2(5.4-5.3)]/5.4 = 0.296
Simulations were applied to management in the late 1950s for queuing and stock control issues. Monte Carlo simulation modelled activities of warehouses, oil depots, and queuing problems like supermarket checkouts. The term “Monte Carlo” comes from the famous gambling city in Monaco, where random numbers from a roulette wheel are similar to random numbers generated by computers in simulations.
The benefits of utilising simulation over queuing theory are as follows:
- It can more effectively manage time-dependent behaviours.
- The mathematics involved in queuing theory is complex and only applicable to certain statistical distributions, whereas the mathematics of simulation is straightforward and adaptable to any statistical distribution.
- In certain scenarios, it is nearly impossible to construct the equations required by queuing theory, such as those involving queue switching or queue-dependent work rates.
- Simulation is significantly easier for managers to comprehend and utilize compared to queuing theory.
However, one drawback of simulation is the challenge in finding optimal solutions. Unlike linear programming, which employs algorithms to automatically determine optimal solutions, optimization in simulation requires a trial-and-error approach:
- Implement a change,
- Run the simulation program to assess whether an improvement has been achieved,
- Repeat the process as necessary.
Once we have the model, we can use it to:
- Understand the current system. For instance, identify factors causing delays in factory production.
- Explore changes to improve the existing system. Examples include adding machines, speeding them up, or reducing idle time through better maintenance.
- Design a new system from scratch. For example, optimizing resource levels and placement in an airport passenger terminal to meet statistical requirements at minimum cost.
Little’s Law
In place of developing complex mathematical models with a number of variables to consider, businesses can use simpler, more flexible models such as Little’s law.
Little’s law is simply written as L=λW, where L represents the average number of items in a system, λ denotes the average arrival rate of items, and W indicates the average wait time of an item.
Little’s law states that the average length of a queue in steady state is given by the product of the rate at which customers enter the line and the average amount of time they spend in line.
Despite the perceived simplicity of the formula, Little’s law is powerful because the relationship is not dependent on the distribution of arriving customers or service time. Little’s law holds for any service order or queuing principle.
An application of Little’s law can provide a simplified analysis of complex systems by explicitly examining just queue length and wait time as opposed to considering a number of other factors. Despite the benefits of Little’s law, such a highly flexible model is not often used in queuing theory because the analysis is perhaps too simplified to provide an accurate representation of queuing models.
Little’s law may still function well in in overviewing ongoing operations. Although Little’s law relates three distinct average statistics, each value is a clear measure of the effectiveness of a process. Any fault in the system would likely be manifested in at least one of the averages.
Little’s Law connects the capacity of a queuing system, the average time spent in the system, and the average arrival rate into the system without knowing any other features of the queue. The formula is quite simple and is written as follows:

or transformed to solve for the other two variables so that:


Where:
- L is the average number of customers in the system
- λ (lambda) is the average arrival rate into the system
- W is the average amount of time spent in the system
Change management processes like Lean and Kanban wouldn’t exist without the Little’s Law queuing models. They’re critical for business applications, simply stated Little’s Law can be written as follows:

Read more about Little’s Law here
A small number of Queuing Theory and Operations Research references
Anderson, David R et al. Quantitative Methods for Business. 12th ed. Mason, Ohio: South-Western, 2012. Print.
Beasley, J E. “OR-Notes.” Brunel University London Personal Pages. N.p., 2005. Print.
Byrd Jack, Jr. “The Value of Queueing Theory.” Interfaces 8.3 (1978): 22–26. Web.
Cope, Robert F, III, Rachelle F Cope, and Harold E Davis. “Disney’s Virtual Queues : A Strategic.” Journal of Business & Economics Research 6.10 (2008): 13–20. Print.
Cope, Rachelle F et al. “Innovative Knowledge Management At Disney : Human Capital And Queuing Solutions For Services.” Journal of Service Science 4.1 (2011): 13–20. Print.
De Lange, Robert, Ilya Samoilovich, and Bo Van Der Rhee. “Virtual Queuing at Airport Security Lanes.” European Journal of Operational Research 225.1 (2013): 153–165. Web.
Dickson, Duncan, Robert C. Ford, and Bruce Laval. “Managing Real and Virtual Waits in Hospitality and Service Organizations.” Cornell Hotel and Restaurant Administration Quarterly 46.1 (2005): 52–68. Web.
Dudin, Alexander et al. “Multi-Server Queueing System with a Generalized Phase-Type Service Time Distribution as a Model of Call Center with a Call-Back Option.” Annals of Operations Research 239.2 (2016): 401–428. Web.
Katz, K, B Larson, and R Larson. “Prescription for the Waiting in Line Blues.” Sloan Management Review Winter.October (1991): 44–53. Print.
Kimes, Sheryl S E. “The Role of Technology in Restaurant Revenue Management.” Cornell
Hospitality Quarterly 49.3 (2008): 297–309. Web.32
Little, John D. C. “Little’s Law as Viewed on Its 50th Anniversary.” Operations Research 59.3 (2011): 536–549. Web.
Nelson, Emily. “The Art of Queueing up at Disneyland.” Journal of Tourism History 8.1 (2016): 47–56. Web.
Niles, Robert. “The Economics of How Disney Got from E Tickets to Paid Fastpasses.” Theme Park Insider. N.p., 2018. Web.
Simchi-Levi, David, and Michael A. Trick. “Introduction to ‘Little’s Law as Viewed on Its 50th Anniversary.’” Operations Research 59.3 (2011): 535–535. Web.
Taha, Hamdy A. “Queueing Theory in Practice.” Interfaces 11.1 (1981): 43–49. Web.
