0% found this document useful (0 votes)

117 views40 pages

Lecture 1: An Introduction To CUDA: Mike Giles

The document provides an overview of CUDA hardware and software. It describes the hardware view of GPUs with multiple cores and device memory. It explains the software view of a CUDA program with a host CPU code and kernel code that runs on the GPU. It also covers CUDA programming components and the basics of launching kernels on the GPU.

Uploaded by

sdancer75

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

117 views40 pages

Lecture 1: An Introduction To CUDA: Mike Giles

Uploaded by

sdancer75

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 40

Lecture 1: an introduction to CUDA

Mike Giles
[email protected]

Oxford University Mathematical Institute

Oxford e-Research Centre

Lecture 1 – p. 1
Overview
hardware view
software view
CUDA programming

Lecture 1 – p. 2
Hardware view
At the top-level, a PCIe graphics card with a many-core
GPU and high-speed graphics “device” memory sits inside
a standard PC/server with one or two multicore CPUs:

GDDR5
DDR4 or HBM

motherboard graphics card

Lecture 1 – p. 3
Hardware view
Currently, 4 generations of hardware cards in use:
Kepler (compute capability 3.x):
first released in 2012, including HPC cards with
excellent DP
our practicals will use K40s and K80s

Maxwell (compute capability 5.x):

first released in 2014; only gaming cards, so poor DP

Pascal (compute capability 6.x):

first released in 2016
many gaming cards and several HPC cards in Oxford

Volta (compute capability 7.x):

first released in 2018; only HPC cards so far Lecture 1 – p. 4
Hardware view
The Pascal generation has cards for both gaming/VR and
HPC

Consumer graphics cards (GeForce):

GTX 1060: 1280 cores, 6GB (£230)
GTX 1070: 1920 cores, 8GB (£380)
GTX 1080: 2560 cores, 8GB (£480)
GTX 1080 Ti: 3584 cores, 11GB (£650)

HPC (Tesla):
P100 (PCIe): 3584 cores, 12GB HBM2 (£5k)
P100 (PCIe): 3584 cores, 16GB HBM2 (£6k)
P100 (NVlink): 3584 cores, 16GB HBM2 (£8k?) Lecture 1 – p. 5
Hardware view
building block is a “streaming multiprocessor” (SM):
128 cores (64 in P100) and 64k registers
96KB (64KB in P100) of shared memory
48KB (24KB in P100) L1 cache
8-16KB (?) cache for constants
up to 2K threads per SM
different chips have different numbers of these SMs:
product SMs bandwidth memory power
GTX 1060 10 192 GB/s 6 GB 120W
GTX 1070 16 256 GB/s 8 GB 150W
GTX 1080 20 320 GB/s 8 GB 180W
GTX Titan X 28 480 GB/s 12 GB 250W
P100 56 720 GB/s 16 GB HBM2 300W
Lecture 1 – p. 6
Hardware View
Pascal GPU ✟✟
✟ ✟
✟ ✏✏
✏
✟✟✏✏

SM SM SM SM

❏ ❅
❏ ❅
L2 cache ❏ ❅
❏ ❅
❏ ❅
❏ shared memory
SM SM SM SM ❏
❏ L1 cache
❏
❏

Lecture 1 – p. 7
Hardware view
There were multiple products in the Kepler generation

Consumer graphics cards (GeForce):

GTX Titan Black: 2880 cores, 6GB
GTX Titan Z: 2×2880 cores, 2×6GB

HPC cards (Tesla):

K20: 2496 cores, 5GB
K40: 2880 cores, 12GB
K80: 2×2496 cores, 2×12GB

Lecture 1 – p. 8
Hardware view
building block is a “streaming multiprocessor” (SM):
192 cores and 64k registers
64KB of shared memory / L1 cache
8KB cache for constants
48KB texture cache for read-only arrays
up to 2K threads per SM

different chips have different numbers of these SMs:

product SMs bandwidth memory power
GTX Titan Z 2×15 2×336 GB/s 2×6 GB 375W
K40 15 288 GB/s 12 GB 245W
K80 2×14 2×240 GB/s 2×12 GB 300W

Lecture 1 – p. 9
Hardware View
Kepler GPU ✟✟
✟ ✟
✟ ✏✏
✏
✟✟✏✏

SM SM SM SM

❏ ❅
❏ ❅
L2 cache ❏ ❅
❏ ❅
❏ ❅
❏
SM SM SM SM ❏ L1 cache /
❏ shared memory
❏
❏

Lecture 1 – p. 10
Multithreading
Key hardware feature is that the cores in a SM are SIMT
(Single Instruction Multiple Threads) cores:
groups of 32 cores execute the same instructions
simultaneously, but with different data
similar to vector computing on CRAY supercomputers
32 threads all doing the same thing at the same time
natural for graphics processing and much scientific
computing
SIMT is also a natural choice for many-core chips to
simplify each core

Lecture 1 – p. 11
Multithreading
Lots of active threads is the key to high performance:
no “context switching”; each thread has its own
registers, which limits the number of active threads
threads on each SM execute in groups of 32 called
“warps” – execution alternates between “active” warps,
with warps becoming temporarily “inactive” when
waiting for data

Lecture 1 – p. 12
Multithreading
originally, each thread completed one operation before
the next started to avoid complexity of pipeline overlaps
✲
✲ ✲ time
✲1 2345
✲ ✲
✲1 2345
✲ ✲
✲1 2345

however, NVIDIA have now relaxed this, so each thread

can have multiple independent instructions overlapping

memory access from device memory has a delay of

200-400 cycles; with 40 active warps this is equivalent
to 5-10 operations, so enough to hide the latency?
Lecture 1 – p. 13
Software view
At the top level, we have a master process which runs on
the CPU and performs the following steps:
1. initialises card
2. allocates memory in host and on device
3. copies data from host to device memory
4. launches multiple instances of execution “kernel” on
device
5. copies data from device memory to host
6. repeats 3-5 as needed
7. de-allocates all memory and terminates

Lecture 1 – p. 14
Software view
At a lower level, within the GPU:
each instance of the execution kernel executes on a SM
if the number of instances exceeds the number of SMs,
then more than one will run at a time on each SM if
there are enough registers and shared memory, and the
others will wait in a queue and execute later
all threads within one instance can access local shared
memory but can’t see what the other instances are
doing (even if they are on the same SM)
there are no guarantees on the order in which the
instances execute

Lecture 1 – p. 15
CUDA
CUDA (Compute Unified Device Architecture) is NVIDIA’s
program development environment:
based on C/C++ with some extensions
FORTRAN support provided by compiler from PGI
(owned by NVIDIA) and also in IBM XL compiler
lots of example code and good documentation
– fairly short learning curve for those with experience of
OpenMP and MPI programming
large user community on NVIDIA forums

Lecture 1 – p. 16
CUDA Components
Installing CUDA on a system, there are 3 components:
driver
low-level software that controls the graphics card
toolkit
nvcc CUDA compiler
Nsight IDE plugin for Eclipse or Visual Studio
profiling and debugging tools
several libraries
SDK
lots of demonstration examples
some error-checking utilities
not officially supported by NVIDIA
almost no documentation Lecture 1 – p. 17
CUDA programming
Already explained that a CUDA program has two pieces:
host code on the CPU which interfaces to the GPU
kernel code which runs on the GPU

At the host level, there is a choice of 2 APIs

(Application Programming Interfaces):
runtime
simpler, more convenient
driver
much more verbose, more flexible (e.g. allows
run-time compilation), closer to OpenCL

We will only use the runtime API in this course, and that is
all I use in my own research.
Lecture 1 – p. 18
CUDA programming
At the host code level, there are library routines for:
memory allocation on graphics card
data transfer to/from device memory
constants
ordinary data
error-checking
timing

There is also a special syntax for launching multiple

instances of the kernel process on the GPU.

Lecture 1 – p. 19
CUDA programming
In its simplest form it looks like:
kernel_routine<<<gridDim, blockDim>>>(args);

gridDim is the number of instances of the kernel

(the “grid” size)
blockDim is the number of threads within each
instance
(the “block” size)
args is a limited number of arguments, usually mainly
pointers to arrays in graphics memory, and some
constants which get copied by value

The more general form allows gridDim and blockDim to

be 2D or 3D to simplify application programs
Lecture 1 – p. 20
CUDA programming
At the lower level, when one instance of the kernel is started
on a SM it is executed by a number of threads,
each of which knows about:
some variables passed as arguments
pointers to arrays in device memory (also arguments)
global constants in device memory
shared memory and private registers/local variables
some special variables:
gridDim size (or dimensions) of grid of blocks
blockDim size (or dimensions) of each block
blockIdx index (or 2D/3D indices) of block
threadIdx index (or 2D/3D indices) of thread
warpSize always 32 so far, but could change Lecture 1 – p. 21
CUDA programming
1D grid with 4 blocks, each with 64 threads:

gridDim = 4
blockDim = 64
blockIdx ranges from 0 to 3
threadIdx ranges from 0 to 63

blockIdx.x=1, threadIdx.x=44

❄
r

Lecture 1 – p. 22
CUDA programming
The kernel code looks fairly normal once you get used to
two things:
code is written from the point of view of a single thread
quite different to OpenMP multithreading
similar to MPI, where you use the MPI “rank” to
identify the MPI process
all local variables are private to that thread
need to think about where each variable lives (more on
this in the next lecture)
any operation involving data in the device memory
forces its transfer to/from registers in the GPU
often better to copy the value into a local register
variable
Lecture 1 – p. 23
Host code
int main(int argc, char **argv) {
float *h_x, *d_x; // h=host, d=device
int nblocks=2, nthreads=8, nsize=2*8;

h_x = (float )malloc(nsizesizeof(float));

cudaMalloc((void **)&d_x,nsize*sizeof(float));

my_first_kernel<<<nblocks,nthreads>>>(d_x);

cudaMemcpy(h_x,d_x,nsize*sizeof(float),
cudaMemcpyDeviceToHost);

for (int n=0; n<nsize; n++)

printf(" n, x = %d %f \n",n,h_x[n]);

cudaFree(d_x); free(h_x);
} Lecture 1 – p. 24
Kernel code
#include <helper_cuda.h>

global void my_first_kernel(float *x)

{
int tid = threadIdx.x + blockDim.x*blockIdx.x;

x[tid] = (float) threadIdx.x;

}

global identifier says it’s a kernel function

each thread sets one element of x array
within each block of threads, threadIdx.x ranges
from 0 to blockDim.x-1, so each thread has a unique
value for tid

Lecture 1 – p. 25
CUDA programming
Suppose we have 1000 blocks, and each one has 128
threads – how does it get executed?

On Kepler hardware, would probably get 8-12 blocks

running at the same time on each SM, and each block
has 4 warps =⇒ 32-48 warps running on each SM

Each clock tick, SM warp scheduler decides which warps

to execute next, choosing from those not waiting for
data coming from device memory (memory latency)
completion of earlier instructions (pipeline delay)

Programmer doesn’t have to worry about this level of detail,

just make sure there are lots of threads / warps
Lecture 1 – p. 26
CUDA programming

Queue of waiting blocks:

Multiple blocks running on each SM:

❄ ❄ ❄ ❄

SM SM SM SM

Lecture 1 – p. 27
CUDA programming
In this simple case, we had a 1D grid of blocks, and a 1D
set of threads within each block.

If we want to use a 2D set of threads, then

blockDim.x, blockDim.y give the dimensions, and
threadIdx.x, threadIdx.y give the thread indices

and to launch the kernel we would use something like

dim3 nthreads(16,4);
my_new_kernel<<<nblocks,nthreads>>>(d_x);
where dim3 is a special CUDA datatype with 3 components
.x,.y,.z each initialised to 1.

Lecture 1 – p. 28
CUDA programming
A similar approach is used for 3D threads and 2D / 3D grids;
can be very useful in 2D / 3D finite difference applications.

How do 2D / 3D threads get divided into warps?

1D thread ID defined by
threadIdx.x +
threadIdx.y * blockDim.x +
threadIdx.z * blockDim.x * blockDim.y
and this is then broken up into warps of size 32.

Lecture 1 – p. 29
Practical 1
start from code shown above (but with comments)
learn how to compile / run code within Nsight IDE
(integrated into Visual Studio for Windows,
or Eclipse for Linux)
test error-checking and printing from kernel functions
modify code to add two vectors together (including
sending them over from the host to the device)
if time permits, look at CUDA SDK examples

Lecture 1 – p. 30
Practical 1
Things to note:
memory allocation
cudaMalloc((void **)&d x, nbytes);

data copying
cudaMemcpy(h x,d x,nbytes,
cudaMemcpyDeviceToHost);
reminder: prefix h and d to distinguish between
arrays on the host and on the device is not mandatory,
just helpful labelling
kernel routine is declared by global prefix, and is
written from point of view of a single thread

Lecture 1 – p. 31
Practical 1
Second version of the code is very similar to first, but uses
an SDK header file for various safety checks – gives useful
feedback in the event of errors.

check for error return codes:

checkCudaErrors( ... );

check for kernel failure messages:

getLastCudaError( ... );

Lecture 1 – p. 32
Practical 1
One thing to experiment with is the use of printf within
a CUDA kernel function:
essentially the same as standard printf; minor
difference in integer return code
each thread generates its own output; use conditional
code if you want output from only one thread
output goes into an output buffer which is transferred
to the host and printed later (possibly much later?)
buffer has limited size (1MB by default), so could lose
some output if there’s too much
need to use either cudaDeviceSynchronize(); or
cudaDeviceReset(); at the end of the main code to
make sure the buffer is flushed before termination
Lecture 1 – p. 33
Practical 1
The practical also has a third version of the code which
uses “managed memory” based on Unified Memory.

In this version
there is only one array / pointer, not one for CPU and
another for GPU
the programmer is not responsible for moving the data
to/from the GPU
everything is handled automatically by the CUDA
run-time system

Lecture 1 – p. 34
Practical 1
This leads to simpler code, but it’s important to understand
what is happening because it may hurt performance:

if the CPU initialises an array x, and then a kernel uses

it, this forces a copy from CPU to GPU
if the GPU modifies x and the CPU later tries to read
from it, that triggers a copy back from GPU to CPU

Personally, I prefer to keep complete control over data

movement, so that I know what is happening and I can
maximise performance.

Lecture 1 – p. 35
ARCUS-B cluster

external network

arcus-b

gnode1101 gnode1102 gnode1103 gnode1104 gnode1105

G G G G G G G G G G

arcus-b.arc.ox.ac.uk is the head node

the GPU compute nodes have two K80 cards with a
total of 4 GPUs, numbered 0 – 3
read the Arcus notes before starting the practical Lecture 1 – p. 36
Key reading
CUDA Programming Guide, version 8.0:
Chapter 1: Introduction
Chapter 2: Programming Model
Section 5.4: performance of different GPUs
Appendix A: CUDA-enabled GPUs
Appendix B, sections B.1 – B.4: C language extensions
Appendix B, section B.17: printf output
Appendix G, section G.1: features of different GPUs

Wikipedia (clearest overview of NVIDIA products):

en.wikipedia.org/wiki/Nvidia Tesla
en.wikipedia.org/wiki/GeForce 10 series
Lecture 1 – p. 37
Nsight
General view:

Lecture 1 – p. 38
Nsight
Importing the practicals: select General – Existing Projects

Lecture 1 – p. 39
Nsight

Lecture 1 – p. 40

STS Advance User Manual
No ratings yet
STS Advance User Manual
89 pages
Sega Saturn Architecture: Architecture of Consoles: A Practical Analysis, #5
From Everand
Sega Saturn Architecture: Architecture of Consoles: A Practical Analysis, #5
Rodrigo Copetti
No ratings yet
Lecture 1: An Introduction To CUDA: Mike Giles
No ratings yet
Lecture 1: An Introduction To CUDA: Mike Giles
247 pages
Lec 1
No ratings yet
Lec 1
27 pages
Lecture 2
No ratings yet
Lecture 2
77 pages
лк CUDA - 1 PDCn
No ratings yet
лк CUDA - 1 PDCn
31 pages
CUDA Tutorial
No ratings yet
CUDA Tutorial
50 pages
Lecture - 01 - CUDA Programming
No ratings yet
Lecture - 01 - CUDA Programming
52 pages
CSE Lec4 Cuda
No ratings yet
CSE Lec4 Cuda
91 pages
1 Cuda
100% (1)
1 Cuda
173 pages
Programming Gpus With Cuda: John Mellor-Crummey
No ratings yet
Programming Gpus With Cuda: John Mellor-Crummey
42 pages
GPGPU Programming With CUDA: Leandro Avila - University of Northern Iowa
No ratings yet
GPGPU Programming With CUDA: Leandro Avila - University of Northern Iowa
29 pages
Chapter 8
No ratings yet
Chapter 8
58 pages
Introduction To Programming Massively Parallel Graphics Processors
No ratings yet
Introduction To Programming Massively Parallel Graphics Processors
84 pages
Chapter7 GPU
No ratings yet
Chapter7 GPU
45 pages
Lecture2 Cuda Basic 2010
No ratings yet
Lecture2 Cuda Basic 2010
44 pages
Gpu Cuda
No ratings yet
Gpu Cuda
204 pages
CUDA
No ratings yet
CUDA
33 pages
Topic GPU1
No ratings yet
Topic GPU1
32 pages
ECE 498AL The CUDA Programming Model
No ratings yet
ECE 498AL The CUDA Programming Model
37 pages
CUDA Programming On Nvidia Gpus: Mike Giles
No ratings yet
CUDA Programming On Nvidia Gpus: Mike Giles
21 pages
0 Gpu Computing I Give It
No ratings yet
0 Gpu Computing I Give It
57 pages
GPU Architecture Ebook
No ratings yet
GPU Architecture Ebook
67 pages
High Performance Computing On Gpu
No ratings yet
High Performance Computing On Gpu
37 pages
002 - Introduction To CUDA Programming - 1
No ratings yet
002 - Introduction To CUDA Programming - 1
54 pages
Lecture 12 GPU Programming
No ratings yet
Lecture 12 GPU Programming
65 pages
CUDA Introduction
No ratings yet
CUDA Introduction
39 pages
CUDA
No ratings yet
CUDA
18 pages
CUDA Introduction Mod
No ratings yet
CUDA Introduction Mod
50 pages
Lecture12 GPUArchCUDA02-CUDAMem
No ratings yet
Lecture12 GPUArchCUDA02-CUDAMem
67 pages
cs179 2016 Lec13
No ratings yet
cs179 2016 Lec13
30 pages
CSED405 Lec2-CUDA Overview - 240916 - 131108
No ratings yet
CSED405 Lec2-CUDA Overview - 240916 - 131108
52 pages
Gpu History and Cuda Programming Basics
No ratings yet
Gpu History and Cuda Programming Basics
44 pages
GPU Programming: Dr. Florian Ferreira
No ratings yet
GPU Programming: Dr. Florian Ferreira
101 pages
CUDA Compute Unified Device Architecture
No ratings yet
CUDA Compute Unified Device Architecture
26 pages
Intro GPUs
No ratings yet
Intro GPUs
36 pages
Gpu1 - GPU Introduction
No ratings yet
Gpu1 - GPU Introduction
20 pages
CUDA Programming
No ratings yet
CUDA Programming
35 pages
GPU Basics
No ratings yet
GPU Basics
93 pages
GPU Programming: CUDA
No ratings yet
GPU Programming: CUDA
29 pages
Lec 2 PDC
No ratings yet
Lec 2 PDC
31 pages
Lecture GPUArchCUDA01
No ratings yet
Lecture GPUArchCUDA01
57 pages
GPU Programming Slides 2
No ratings yet
GPU Programming Slides 2
37 pages
Cuda 1
No ratings yet
Cuda 1
45 pages
L 3 GPU
No ratings yet
L 3 GPU
33 pages
A Beginner'S Guide To Programming Gpus With Cuda: Mike Peardon
No ratings yet
A Beginner'S Guide To Programming Gpus With Cuda: Mike Peardon
21 pages
Lecture3 Fundamentals of CUDA (Part1) - 2025
No ratings yet
Lecture3 Fundamentals of CUDA (Part1) - 2025
52 pages
GPU Cluster4
No ratings yet
GPU Cluster4
31 pages
Cuda
No ratings yet
Cuda
25 pages
27th Aug - Introduction To GPGPU - Part 1
No ratings yet
27th Aug - Introduction To GPGPU - Part 1
32 pages
Unit 6 Chapter 1 Parallel Programming Tools Cuda - Programming
No ratings yet
Unit 6 Chapter 1 Parallel Programming Tools Cuda - Programming
28 pages
Threads
No ratings yet
Threads
54 pages
Kirk+Hwu GPU
No ratings yet
Kirk+Hwu GPU
92 pages
Chapter 5 - General Purpose PGPU, CUDA
No ratings yet
Chapter 5 - General Purpose PGPU, CUDA
70 pages
Parallel Processing With Cuda
No ratings yet
Parallel Processing With Cuda
25 pages
DS1822 - Parallel Computing-Unit3
No ratings yet
DS1822 - Parallel Computing-Unit3
17 pages
21.L18 Intro To GPU and CUDA C
No ratings yet
21.L18 Intro To GPU and CUDA C
89 pages
8 Cud A 1
No ratings yet
8 Cud A 1
38 pages
Nintendo 64 Architecture: Architecture of Consoles: A Practical Analysis, #8
From Everand
Nintendo 64 Architecture: Architecture of Consoles: A Practical Analysis, #8
Rodrigo Copetti
No ratings yet
Master System Architecture: Architecture of Consoles: A Practical Analysis, #15
From Everand
Master System Architecture: Architecture of Consoles: A Practical Analysis, #15
Rodrigo Copetti
2/5 (1)
Mega Drive Architecture: Architecture of Consoles: A Practical Analysis, #3
From Everand
Mega Drive Architecture: Architecture of Consoles: A Practical Analysis, #3
Rodrigo Copetti
No ratings yet
How Do I Disable X at Boot Time So That The System Boots in Text Mode
No ratings yet
How Do I Disable X at Boot Time So That The System Boots in Text Mode
11 pages
Concurrent Kernel in OpenCL
No ratings yet
Concurrent Kernel in OpenCL
1 page
(INFO) ANDROID DEVICE PARTITIONS and FILESYSTEMS - XDA Developers Forums
No ratings yet
(INFO) ANDROID DEVICE PARTITIONS and FILESYSTEMS - XDA Developers Forums
12 pages
Opencl 2.0 Features: Benjamin Coquelle MAY 2015
No ratings yet
Opencl 2.0 Features: Benjamin Coquelle MAY 2015
40 pages
Redirect HTTP To HTTPS in Nginx - Linuxize
No ratings yet
Redirect HTTP To HTTPS in Nginx - Linuxize
7 pages
Android Partitions Explained - Boot, System, Recovery, Data, Cache & Misc
No ratings yet
Android Partitions Explained - Boot, System, Recovery, Data, Cache & Misc
6 pages
Token Based Authentication Made Easy - Auth0
100% (1)
Token Based Authentication Made Easy - Auth0
10 pages
How To Enable ES6 (And Beyond) Syntax With Node and Express
No ratings yet
How To Enable ES6 (And Beyond) Syntax With Node and Express
20 pages
Performance Microsoft - TypeScript Wiki
No ratings yet
Performance Microsoft - TypeScript Wiki
11 pages
Inside The Intel and Creative Assembly Collaboration: White Paper
No ratings yet
Inside The Intel and Creative Assembly Collaboration: White Paper
10 pages
Redis Command Line To View Chinese Without Scrambling
No ratings yet
Redis Command Line To View Chinese Without Scrambling
3 pages
Pre-73 DLX: Vintage Style Pre Amplifier
No ratings yet
Pre-73 DLX: Vintage Style Pre Amplifier
2 pages
How To Compute The PSNR (Peak Signal-To-Noise Ratio)
No ratings yet
How To Compute The PSNR (Peak Signal-To-Noise Ratio)
3 pages
8 Function Christmas Light Circuit - Homemade Circuit Projects
No ratings yet
8 Function Christmas Light Circuit - Homemade Circuit Projects
5 pages
G5ca 1a Relay
No ratings yet
G5ca 1a Relay
4 pages
The Use of Information Technology in The Universit
No ratings yet
The Use of Information Technology in The Universit
9 pages
A Journey Through The CPU Pipeline
No ratings yet
A Journey Through The CPU Pipeline
20 pages
Manuale Ecoloader (Eng) - Rev2014
No ratings yet
Manuale Ecoloader (Eng) - Rev2014
16 pages
F3294 Phe840m
No ratings yet
F3294 Phe840m
2 pages
Unit 1 - IMED
No ratings yet
Unit 1 - IMED
37 pages
CAPE Info Tech Unit 2 P1 2011
No ratings yet
CAPE Info Tech Unit 2 P1 2011
9 pages
CPSQ Report
No ratings yet
CPSQ Report
11 pages
8051 Lab Manual
No ratings yet
8051 Lab Manual
55 pages
Date-Sheet, B.tech Dec-2010 & Jan-2011
No ratings yet
Date-Sheet, B.tech Dec-2010 & Jan-2011
15 pages
Vintron PoE Switch
No ratings yet
Vintron PoE Switch
11 pages
Progress Draft Propossal Deffense 1
No ratings yet
Progress Draft Propossal Deffense 1
14 pages
Aruba AOS-CX Simulator - GNS3-VM Deployment Guide
No ratings yet
Aruba AOS-CX Simulator - GNS3-VM Deployment Guide
11 pages
417 Competency Based Question Artificial Intelligence Chap-1 (2024-25)
No ratings yet
417 Competency Based Question Artificial Intelligence Chap-1 (2024-25)
4 pages
Intro To PLC
100% (1)
Intro To PLC
2 pages
Excel Int 2days
No ratings yet
Excel Int 2days
1 page
CSS - Info Sheet 3.2-3 - Confirm Network Services To Be Configured
No ratings yet
CSS - Info Sheet 3.2-3 - Confirm Network Services To Be Configured
26 pages
GFWLIVESetup Log Verbose
No ratings yet
GFWLIVESetup Log Verbose
1 page
Python Model Paper-01
No ratings yet
Python Model Paper-01
28 pages
11.2 Analyze and Store Logs
No ratings yet
11.2 Analyze and Store Logs
6 pages
Built in Data Type
No ratings yet
Built in Data Type
19 pages
DS 3E2528B Gigabit Industrial Switch
No ratings yet
DS 3E2528B Gigabit Industrial Switch
6 pages
Coding Form Penghitungan
No ratings yet
Coding Form Penghitungan
4 pages
DDSM User Manual
No ratings yet
DDSM User Manual
10 pages
It Assignment - 1: Analog Signal Digital Signal
No ratings yet
It Assignment - 1: Analog Signal Digital Signal
6 pages
04 EasyIO FS Series Firmware Upgrade v1.2
No ratings yet
04 EasyIO FS Series Firmware Upgrade v1.2
13 pages
Introduction To Advanced Semiconductor Memories
No ratings yet
Introduction To Advanced Semiconductor Memories
18 pages
ITAPBUS Notes
No ratings yet
ITAPBUS Notes
5 pages
DS Lecture-1
No ratings yet
DS Lecture-1
93 pages
VMware l2p 1V0-601 v2015-08-10 by Far 110q
No ratings yet
VMware l2p 1V0-601 v2015-08-10 by Far 110q
34 pages
Chapter 5 - Pointers
No ratings yet
Chapter 5 - Pointers
63 pages
RN21009 CoLOS Product Suite 6.4
No ratings yet
RN21009 CoLOS Product Suite 6.4
12 pages
Ubiquiti UF OLT 4 Quick Start Guide A1
No ratings yet
Ubiquiti UF OLT 4 Quick Start Guide A1
14 pages
DWIN Panel and Siemens PLC Hardware and Software Connect
No ratings yet
DWIN Panel and Siemens PLC Hardware and Software Connect
7 pages

Lecture 1: An Introduction To CUDA: Mike Giles

Uploaded by

Lecture 1: An Introduction To CUDA: Mike Giles

Uploaded by

Lecture 1: an introduction to CUDA

Oxford University Mathematical Institute

motherboard graphics card

Maxwell (compute capability 5.x):

Pascal (compute capability 6.x):

Volta (compute capability 7.x):

Consumer graphics cards (GeForce):

Consumer graphics cards (GeForce):

HPC cards (Tesla):

different chips have different numbers of these SMs:

however, NVIDIA have now relaxed this, so each thread

memory access from device memory has a delay of

At the host level, there is a choice of 2 APIs

There is also a special syntax for launching multiple

gridDim is the number of instances of the kernel

The more general form allows gridDim and blockDim to

h_x = (float *)malloc(nsize*sizeof(float));

for (int n=0; n<nsize; n++)

__global__ void my_first_kernel(float *x)

x[tid] = (float) threadIdx.x;

global identifier says it’s a kernel function

On Kepler hardware, would probably get 8-12 blocks

Each clock tick, SM warp scheduler decides which warps

Programmer doesn’t have to worry about this level of detail,

Queue of waiting blocks:

Multiple blocks running on each SM:

If we want to use a 2D set of threads, then

and to launch the kernel we would use something like

How do 2D / 3D threads get divided into warps?

check for error return codes:

check for kernel failure messages:

if the CPU initialises an array x, and then a kernel uses

Personally, I prefer to keep complete control over data

gnode1101 gnode1102 gnode1103 gnode1104 gnode1105

arcus-b.arc.ox.ac.uk is the head node

Wikipedia (clearest overview of NVIDIA products):

You might also like

h_x = (float )malloc(nsizesizeof(float));

global void my_first_kernel(float *x)