0% found this document useful (0 votes)

42 views7 pages

Shared-Memory Multiprocessors - Symmetric Multiprocessing Hardware

The document discusses symmetric multiprocessing hardware architectures. It describes bus and crossbar architectures that connect multiple processors to shared memory. A bus is simpler but can become a bottleneck, while a crossbar eliminates this bottleneck but has higher costs. Cache memory helps mitigate bandwidth issues by reducing main memory traffic. The document also discusses cache coherency protocols that ensure processors see consistent values when variables are shared between caches.

Uploaded by

Silvio Dresser

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

42 views7 pages

Shared-Memory Multiprocessors - Symmetric Multiprocessing Hardware

Uploaded by

Silvio Dresser

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 7

Connexions module: m32794

Shared-Memory Multiprocessors Symmetric Multiprocessing

Hardware

Charles Severance
This work is produced by The Connexions Project and licensed under the
Creative Commons Attribution License
In Figure 1 (A shared-memory multiprocessor), we viewed an ideal shared-memory multiprocessor. In
this section, we look in more detail at how such a system is actually constructed. The primary advantage
of these systems is the ability for any CPU to access all of the memory and peripherals. Furthermore, the
systems need a facility for deciding among themselves who has access to what, and when, which means there
will have to be hardware support for arbitration.
symmetric multiprocessing are

buses

and

crossbars.

The two most common architectural underpinnings for

The bus is the simplest of the two approaches. Figure 2

(A typical bus architecture) shows processors connected using a bus. A bus can be thought of as a set of
parallel wires connecting the components of the computer (CPU, memory, and peripheral controllers), a set
of protocols for communication, and some hardware to help carry it out. A bus is less expensive to build,
but because all trac must cross the bus, as the load increases, the bus eventually becomes a performance
bottleneck.
Version

1.2: Feb 24, 2010 4:36 pm US/Central

http://creativecommons.org/licenses/by/3.0/

http://cnx.org/content/m32794/1.2/

Connexions module: m32794

A shared-memory multiprocessor

Figure 1

A typical bus architecture

Figure 2

A crossbar is a hardware approach to eliminate the bottleneck caused by a single bus.

A crossbar is

like several buses running side by side with attachments to each of the modules on the machine CPU,
memory, and peripherals. Any module can get to any other by a path through the crossbar, and multiple

http://cnx.org/content/m32794/1.2/

Connexions module: m32794

paths may be active simultaneously. In the 45 crossbar of Figure 3 (A crossbar), for instance, there can
be four active data transfers in progress at one time. In the diagram it looks like a patchwork of wires, but
there is actually quite a bit of hardware that goes into constructing a crossbar. Not only does the crossbar
connect parties that wish to communicate, but it must also actively arbitrate between two or more CPUs
that want access to the same memory or peripheral. In the event that one module is too popular, it's the
crossbar that decides who gets access and who doesn't. Crossbars have the best performance because there
is no single shared bus. However, they are more expensive to build, and their cost increases as the number
of ports is increased. Because of their cost, crossbars typically are only found at the high end of the price
and performance spectrum.
Whether the system uses a bus or crossbar, there is only so much memory bandwidth to go around; four
or eight processors drawing from one memory system can quickly saturate all available bandwidth. All of
the techniques that improve memory performance (as described in ) also apply here in the design of the
memory subsystems attached to these buses or crossbars.

A crossbar

Figure 3

1 The Eect of Cache

The most common multiprocessing system is made up of commodity processors connected to memory and
peripherals through a bus. Interestingly, the fact that these processors make use of cache somewhat mitigates
the bandwidth bottleneck on a bus-based architecture. By connecting the processor to the cache and viewing
the main memory through the cache, we signicantly reduce the memory trac across the bus.

In this

architecture, most of the memory accesses across the bus take the form of cache line loads and ushes. To
understand why, consider what happens when the cache hit rate is very high. In Figure 4 (High cache hit rate

http://cnx.org/content/m32794/1.2/

Connexions module: m32794

reduces main memory trac), a high cache hit rate eliminates some of the trac that would have otherwise
gone out across the bus or crossbar to main memory. Again, it is the notion of locality of reference that
makes the system work. If you assume that a fair number of the memory references will hit in the cache, the
equivalent attainable main memory bandwidth is more than the bus is actually capable of. This assumption
explains why multiprocessors are designed with less bus bandwidth than the sum of what the CPUs can
consume at once.
Imagine a scenario where two CPUs are accessing dierent areas of memory using unit stride. Both CPUs
access the rst element in a cache line at the same time. The bus arbitrarily allows one CPU access to the
memory. The rst CPU lls a cache line and begins to process the data. The instant the rst CPU has
completed its cache line ll, the cache line ll for the second CPU begins. Once the second cache line ll has
completed, the second CPU begins to process the data in its cache line. If the time to process the data in
a cache line is longer than the time to ll a cache line, the cache line ll for processor two completes before
the next cache line request arrives from processor one. Once the initial conict is resolved, both processors
appear to have conict-free access to memory for the remainder of their unit-stride loops.

High cache hit rate reduces main memory trac

Figure 4

In actuality, on some of the fastest bus-based systems, the memory bus is suciently fast that up to
20 processors can access memory using unit stride with very little conict. If the processors are accessing
memory using non-unit stride, bus and memory bank conict becomes apparent, with fewer processors.
This bus architecture combined with local caches is very popular for general-purpose multiprocessing
loads. The memory reference patterns for database or Internet servers generally consist of a combination of
time periods with a small working set, and time periods that access large data structures using unit stride.
Scientic codes tend to perform more non-unit-stride access than general-purpose codes. For this reason,
the most expensive parallel-processing systems targeted at scientic codes tend to use crossbars connected
to multibanked memory systems.

http://cnx.org/content/m32794/1.2/

Connexions module: m32794

The main memory system is better shielded when a larger cache is used. For this reason, multiprocessors
sometimes incorporate a two-tier cache system, where each processor uses its own small on-chip local cache,
backed up by a larger second board-level cache with as much as 4 MB of memory. Only when neither can
satisfy a memory request, or when data has to be written back to main memory, does a request go out over
the bus or crossbar.

2 Coherency
Now, what happens when one CPU of a multiprocessor running a single program in parallel changes the
value of a variable, and another CPU tries to read it? Where does the value come from? These questions
are interesting because there can be multiple copies of each variable, and some of them can hold old or stale
values.
For illustration, say that you are running a program with a shared variable A. Processor 1 changes the
value of A and Processor 2 goes to read it.

Multiple copies of variable A

Figure 5

In Figure 5 (Multiple copies of variable A), if Processor 1 is keeping

A as a register-resident variable, then

Processor 2 doesn't stand a chance of getting the correct value when it goes to look for it. There is no way
that 2 can know the contents of 1's registers; so assume, at the very least, that Processor 1 writes the new
value back out. Now the question is, where does the new value get stored? Does it remain in Processor 1's
cache? Is it written to main memory? Does it get updated in Processor 2's cache?
Really, we are asking what kind of

cache coherency protocol

see a uniform view of the values in memory.

the vendor uses to assure that all processors

It generally isn't something that the programmer has to

worry about, except that in some cases, it can aect performance. The approaches used in these systems
are similar to those used in single-processor systems with some extensions. The most straight-forward cache
coherency approach is called a

write-through policy

: variables written into cache are simultaneously written

into main memory. As the update takes place, other caches in the system see the main memory reference
being performed. This can be done because all of the caches continuously monitor (also known as

snooping

) the trac on the bus, checking to see if each address is in their cache. If a cache notices that it contains
a copy of the data from the locations being written, it may either

invalidate

its copy of the variable or

obtain new values (depending on the policy). One thing to note is that a write-through cache demands a

http://cnx.org/content/m32794/1.2/

Connexions module: m32794

fair amount of main memory bandwidth since each write goes out over the main memory bus. Furthermore,
successive writes to the same location or bank are subject to the main memory cycle time and can slow the
machine down.
A more sophisticated cache coherency protocol is called

copyback

writeback.

The idea is that you

write values back out to main memory only when the cache housing them needs the space for something else.
Updates of cached data are coordinated between the caches, by the caches, without help from the processor.
Copyback caching also uses hardware that can monitor (snoop) and respond to the memory transactions of
the other caches in the system. The benet of this method over the write-through method is that memory
trac is reduced considerably. Let's walk through it to see how it works.

3 Cache Line States

For this approach to work, each cache must maintain a state for each line in its cache. The possible states
used in the example include:

Modied: This cache line needs to be written back to memory.

Exclusive: There are no other caches that have this cache line.
Shared: There are read-only copies of this line in two or more caches.
Empty/Invalid: This cache line doesn't contain any useful data.
This particular coherency protocol is often called

MESI.

Other cache coherency protocols are more compli-

cated, but these states give you an idea how multiprocessor writeback cache coherency works.
We start where a particular cache line is in memory and in none of the writeback caches on the systems.
The rst cache to ask for data from a particular part of memory completes a normal memory access; the
main memory system returns data from the requested location in response to a cache miss. The associated
cache line is marked

exclusive,

meaning that this is the only cache in the system containing a copy of the

data; it is the owner of the data. If another cache goes to main memory looking for the same thing, the
request is intercepted by the rst cache, and the data is returned from the rst cache not main memory.
Once an interception has occurred and the data is returned, the data is marked

shared in both of the caches.

When a particular line is marked shared, the caches have to treat it dierently than they would if they
were the exclusive owners of the data especially if any of them wants to modify it. In particular, a write to
a shared cache entry is preceded by a broadcast message to all the other caches in the system. It tells them
to invalidate their copies of the data. The one remaining cache line gets marked as

modied

to signal that

it has been changed, and that it must be returned to main memory when the space is needed for something
else.

By these mechanisms, you can maintain cache coherence across the multiprocessor without adding

tremendously to the memory trac.

By the way, even if a variable is not shared, it's possible for copies of it to show up in several caches.
On a symmetric multiprocessor, your program can bounce around from CPU to CPU. If you run for a little
while on this CPU, and then a little while on that, your program will have operated out of separate caches.
That means that there can be several copies of seemingly unshared variables scattered around the machine.
Operating systems often try to minimize how often a process is moved between physical CPUs during context
switches. This is one reason not to overload the available processors in a system.

4 Data Placement
There is one more pitfall regarding shared memory we have so far failed to mention.

It involves data

movement. Although it would be convenient to think of the multiprocessor memory as one big pool, we have
seen that it is actually a carefully crafted system of caches, coherency protocols, and main memory. The
problems come when your application causes lots of data to be traded between the caches. Each reference
that falls out of a given processor's cache (especially those that require an update in another processor's
cache) has to go out on the bus.

http://cnx.org/content/m32794/1.2/

Connexions module: m32794

Often, it's slower to get memory from another processor's cache than from the main memory because of
the protocol and processing overhead involved. Not only do we need to have programs with high locality of
reference and unit stride, we also need to minimize the data that must be moved from one CPU to another.

http://cnx.org/content/m32794/1.2/

Multiprocessor System and Interconnection Networks
No ratings yet
Multiprocessor System and Interconnection Networks
66 pages
CS6801 MULTI CORE ARCHITECTURE AND PROGRAMMING - Watermark
No ratings yet
CS6801 MULTI CORE ARCHITECTURE AND PROGRAMMING - Watermark
96 pages
Carton Packaging Knowledge
88% (8)
Carton Packaging Knowledge
93 pages
Lec13 Multiprocessors
No ratings yet
Lec13 Multiprocessors
69 pages
CH20 COA11e
No ratings yet
CH20 COA11e
40 pages
1.symmetric and Distributed Shared Memory Architectures
79% (19)
1.symmetric and Distributed Shared Memory Architectures
29 pages
Week 5
No ratings yet
Week 5
52 pages
VII. Cache Coherence. Interconnection Networks (1) : March 16, 2009
No ratings yet
VII. Cache Coherence. Interconnection Networks (1) : March 16, 2009
42 pages
05 Multiprocessor
No ratings yet
05 Multiprocessor
54 pages
Distributed Shared Memory
No ratings yet
Distributed Shared Memory
23 pages
Shared Memory Architectures
No ratings yet
Shared Memory Architectures
34 pages
4 TH
No ratings yet
4 TH
84 pages
Unit 5
No ratings yet
Unit 5
89 pages
15CS72ACA Module 4 ACA
No ratings yet
15CS72ACA Module 4 ACA
95 pages
Cache Coherency
No ratings yet
Cache Coherency
19 pages
atII Bks Lec 2021 31 32
No ratings yet
atII Bks Lec 2021 31 32
16 pages
2.radmi 2013 Vol 2 Multiprocessorinterconnectionnetworks
No ratings yet
2.radmi 2013 Vol 2 Multiprocessorinterconnectionnetworks
8 pages
Unit-5 Part-2
No ratings yet
Unit-5 Part-2
22 pages
Unit 6
No ratings yet
Unit 6
36 pages
ch.4 and 5
No ratings yet
ch.4 and 5
41 pages
2ad6a430 1637912349895
No ratings yet
2ad6a430 1637912349895
51 pages
Chapter 5 - Shared Memory Multiprocessor
No ratings yet
Chapter 5 - Shared Memory Multiprocessor
96 pages
MODULE 4 HPC
No ratings yet
MODULE 4 HPC
41 pages
Pipeline
No ratings yet
Pipeline
43 pages
Comporg6 ch12
No ratings yet
Comporg6 ch12
36 pages
VP Interconnection Networks 1
No ratings yet
VP Interconnection Networks 1
18 pages
Module 3
No ratings yet
Module 3
25 pages
Lec 6 SharedArch PDF
No ratings yet
Lec 6 SharedArch PDF
33 pages
Unit 5
No ratings yet
Unit 5
23 pages
Multi-Processor / Parallel Processing
No ratings yet
Multi-Processor / Parallel Processing
12 pages
Parallelism and Multicores
No ratings yet
Parallelism and Multicores
54 pages
Multi-Processor-Parallel Processing PDF
No ratings yet
Multi-Processor-Parallel Processing PDF
12 pages
CO Unit6
No ratings yet
CO Unit6
8 pages
Shared Memory. Distributed Memory. Hybrid Distributed-Shared Memory
No ratings yet
Shared Memory. Distributed Memory. Hybrid Distributed-Shared Memory
22 pages
MCA Operating System and Unix Shell Programming 15
No ratings yet
MCA Operating System and Unix Shell Programming 15
12 pages
Cka Practice Questions
100% (1)
Cka Practice Questions
11 pages
Maintenance of Plastics Processing & Testing Machinery Unit 1
100% (5)
Maintenance of Plastics Processing & Testing Machinery Unit 1
41 pages
Unit VI
No ratings yet
Unit VI
50 pages
Chapter 8 - Parallel Processing
No ratings yet
Chapter 8 - Parallel Processing
50 pages
Shared Memory Architecture
No ratings yet
Shared Memory Architecture
39 pages
Multiprocessors
No ratings yet
Multiprocessors
39 pages
COA Group Assigment
No ratings yet
COA Group Assigment
11 pages
Unit6 - Microprocessor - Final 1
No ratings yet
Unit6 - Microprocessor - Final 1
30 pages
Multi Processors and Thread Level Parallelism
No ratings yet
Multi Processors and Thread Level Parallelism
74 pages
Unit 11
No ratings yet
Unit 11
10 pages
Interconnection Networks: Prof. Varsha Poddar Department of CSE
No ratings yet
Interconnection Networks: Prof. Varsha Poddar Department of CSE
18 pages
COA Assignment
No ratings yet
COA Assignment
21 pages
CA-unit 5-Material-For Reference
No ratings yet
CA-unit 5-Material-For Reference
16 pages
Interconnection Structures
No ratings yet
Interconnection Structures
7 pages
Multiprocessor Architecture and Programming
No ratings yet
Multiprocessor Architecture and Programming
20 pages
A502018463 23825 5 2019 Unit6
No ratings yet
A502018463 23825 5 2019 Unit6
36 pages
Distributed Operating Syst EM: 15SE327E Unit 1
No ratings yet
Distributed Operating Syst EM: 15SE327E Unit 1
49 pages
3 RD
No ratings yet
3 RD
4 pages
Multiprocessors and Multithreading: CS151B/EE M116C Computer Systems Architecture
No ratings yet
Multiprocessors and Multithreading: CS151B/EE M116C Computer Systems Architecture
13 pages
Multi-Processor / Parallel Processing
No ratings yet
Multi-Processor / Parallel Processing
12 pages
07 Multiprocessors MF PDF
No ratings yet
07 Multiprocessors MF PDF
99 pages
MultiProcessors Tanenbaum BP
No ratings yet
MultiProcessors Tanenbaum BP
29 pages
Class 11 CHAPTER 04 by Arslan Saleem
No ratings yet
Class 11 CHAPTER 04 by Arslan Saleem
12 pages
Count Vlad-Jenny Dooley
0% (1)
Count Vlad-Jenny Dooley
35 pages
Within High Fences-Penny Hancock
50% (2)
Within High Fences-Penny Hancock
21 pages
What Is Parallel Computing
No ratings yet
What Is Parallel Computing
9 pages
Multiprocessing: Flynn's Classification (1966)
No ratings yet
Multiprocessing: Flynn's Classification (1966)
8 pages
Multiprocessors and Thread
No ratings yet
Multiprocessors and Thread
4 pages
Method Statement For Installation
No ratings yet
Method Statement For Installation
6 pages
Investor Behaviuor in Volatile Market
No ratings yet
Investor Behaviuor in Volatile Market
66 pages
1 s2.0 S0263224113006519 Main
No ratings yet
1 s2.0 S0263224113006519 Main
11 pages
CPAP-HFNC - Medin - NC3 Ops - Manual Book
No ratings yet
CPAP-HFNC - Medin - NC3 Ops - Manual Book
59 pages
Anthology Poems For Nostalgia
100% (1)
Anthology Poems For Nostalgia
7 pages
AB Salts WKST Key
No ratings yet
AB Salts WKST Key
10 pages
Surface Roughness
No ratings yet
Surface Roughness
8 pages
Sandeep Julakanti - Resume
No ratings yet
Sandeep Julakanti - Resume
9 pages
Chapter 8-Performance Management
No ratings yet
Chapter 8-Performance Management
14 pages
Operating Systems
No ratings yet
Operating Systems
7 pages
How To Use DNA Baser - 2 Minutes Video Tutorial - Url
No ratings yet
How To Use DNA Baser - 2 Minutes Video Tutorial - Url
13 pages
Mitochondrial Disorders Biochemical and Molecular Analysis Methods in Molecular Biology Vol 837 2012th Edition Lee-Jun C. Wong (Editor) Download PDF
100% (2)
Mitochondrial Disorders Biochemical and Molecular Analysis Methods in Molecular Biology Vol 837 2012th Edition Lee-Jun C. Wong (Editor) Download PDF
84 pages
FCE Sample Use of English 1, Twins, Edinburugh, Languages
No ratings yet
FCE Sample Use of English 1, Twins, Edinburugh, Languages
6 pages
Parallel Computing: Charles Koelbel
No ratings yet
Parallel Computing: Charles Koelbel
12 pages
S5 Math Exercise
No ratings yet
S5 Math Exercise
6 pages
23PGHR023 Final Review Ather
No ratings yet
23PGHR023 Final Review Ather
13 pages
Agam
No ratings yet
Agam
12 pages
Group Assignment 6 ICT (XII IPA 5) - 20240118 - 003400 - 0000
No ratings yet
Group Assignment 6 ICT (XII IPA 5) - 20240118 - 003400 - 0000
13 pages
Benchmark Report - Voice Service Optimization For Common State, TP20160728
No ratings yet
Benchmark Report - Voice Service Optimization For Common State, TP20160728
16 pages
A Quantitative Study On The Condom
No ratings yet
A Quantitative Study On The Condom
12 pages
CSC10004: Data Structures and Algorithms
No ratings yet
CSC10004: Data Structures and Algorithms
20 pages
#6 Adding File Upload To A Form
No ratings yet
#6 Adding File Upload To A Form
10 pages
Asia-Pacific Trade Agreement
No ratings yet
Asia-Pacific Trade Agreement
2 pages
The Chevron Way
No ratings yet
The Chevron Way
7 pages
Hipotesis Uji T Kontrol Dan Intervensi
No ratings yet
Hipotesis Uji T Kontrol Dan Intervensi
3 pages
2ND Performance Task in Science
No ratings yet
2ND Performance Task in Science
6 pages
Adjeivos Comparativos y Superativos Teoria y Practica
No ratings yet
Adjeivos Comparativos y Superativos Teoria y Practica
4 pages
TSR Notes
No ratings yet
TSR Notes
6 pages

Shared-Memory Multiprocessors - Symmetric Multiprocessing Hardware

Uploaded by

Shared-Memory Multiprocessors - Symmetric Multiprocessing Hardware

Uploaded by

Connexions module: m32794

Shared-Memory Multiprocessors Symmetric Multiprocessing

The two most common architectural underpinnings for

1.2: Feb 24, 2010 4:36 pm US/Central

Connexions module: m32794

A typical bus architecture

A crossbar is a hardware approach to eliminate the bottleneck caused by a single bus.

Connexions module: m32794

1 The Eect of Cache

Connexions module: m32794

High cache hit rate reduces main memory trac

Connexions module: m32794

Multiple copies of variable A

In Figure 5 (Multiple copies of variable A), if Processor 1 is keeping

A as a register-resident variable, then

cache coherency protocol

see a uniform view of the values in memory.

the vendor uses to assure that all processors

It generally isn't something that the programmer has to

: variables written into cache are simultaneously written

its copy of the variable or

Connexions module: m32794

The idea is that you

3 Cache Line States

Modied: This cache line needs to be written back to memory.

Other cache coherency protocols are more compli-

shared in both of the caches.

tremendously to the memory trac.

Connexions module: m32794

You might also like

1 The Eect of Cache

High cache hit rate reduces main memory trac

see a uniform view of the values in memory.

Modied: This cache line needs to be written back to memory.

tremendously to the memory trac.