I’m happy to announce the release of AMD APP SDK 3.0 supporting OpenCL™ 2.0, the latest compute language API from Khronos. AMD APP SDK 3.0 adds support for Windows® 10 as well as AMD’s latest 6th generation AMD A-series processors (code-named Carrizo), Radeon™ R9 series graphics cards (code-named Fiji) and FirePro™ W8100 and W9100 series graphics cards. OpenCL 2.0 provides several improvements to the programming model that make it much easier to tap into the potential of AMD’s latest GPUs. Notably, Shared Virtual Memory (SVM) allows you to share data-objects that include pointers between the CPU and GPU. This programming construct has long been available on multi-core CPUs, and is now available for AMD GPUs. Sharing memory with pointer-based data structures between GPU and CPU devices greatly simplifies the steps involved in enlisting the GPU for compute acceleration. Also new with OpenCL 2.0 is Device Enqueue (the ability for the GPU device to initiate compute tasks) and support for Generic address space. These features open up a much more powerful programming model for compute kernels.
We gave Gary Demos (one of the pioneers of digital image and video processing technology) an early look at the SDK and here’s what he had to say: “OpenCL 2.0 widens the frontiers of the types of algorithms that can be implemented at high performance. The AMD APP SDK 3.0, together with the OpenCL 2.0 documentation, show how to do things previously impossible on the parallel architecture within GPU’s.”
Sample Code for OpenCL 2.0
AMD APP SDK 3.0 contains a complete set of sample code illustrating how to utilize each of the major new features of OpenCL 2.0. Some of these features are highlighted in the OpenCL 2.0 Demystified blog-series.
There are a few other items also worth mentioning for AMD APP SDK 3.0. We’ve improved the installation process by providing a Web-based installer for Windows users that allows you to only download what you need, but still allows the downloaded package to be distributed locally to the rest of your team. We’ve also updated the OpenCL Programming Guide with many improvements – including full coverage of OpenCL 2.0 features. Check it out – we think it’s the best OpenCL programmers guide available.
To use the AMD APP SDK 3.0, download and install the SDK. Then go to the AMD driver download page and pick the driver you need.
- For AMD APUs and Radeon Graphics, pick the latest AMD Catalyst™ driver for your OS
- For FirePro Graphics, manually select your hardware and OS, and download the driver
Then head to the blogs for more on OpenCL 2.0, or dive into the examples in the SDK, and have fun.
We are always eager for your feedback. Give us kudos, complaints, and suggestions at the AMD OpenCL developer forum. We will listen eagerly to your feedback, and if possible incorporate it in future releases.
Here’s a list of the new and updated samples in the SDK. I first published this back in December when we released the Beta, but it bears repeating here. I’ve included additions since the Beta was released. A quick look at the list will give you an idea of the new power and capability available to you in version 3.0 of the AMD APP SDK.
Samples
| New Samples | ||
|---|---|---|
| Sample | OpenCL™ 2.0 feature | Description |
| SVMBinaryTreeSearch | SVM Coarse Grain | demonstrates the coarse-grain Shared Virtual Memory (SVM) feature of OpenCL 2.0 using a Binary Tree search algorithm |
| SimplePipe | Pipe | demonstrates the Pipe memory object and its APIs |
| PipeProducerConsumerKernels | Pipe | demonstrates the Pipe as a data-sharing FIFO for a producer kernel and a consumer kernel |
| BuiltInScan | New Workgroup Built-in APIs | demonstrates the work group level scan and work group level broadcast features introduced in OpenCL 2.0 using the PrefixSum algorithm |
| ImageBinarization | Image Read and Write | demonstrates using images with read_write qualifier support, which is new in OpenCL 2.0 |
| RecursiveGaussian_ProgramScope | Program Scope Variable | demonstrates Program Scope Variables, a new feature of OpenCL 2.0, using a Recursive Gaussian filter implementation |
| SimpleGenericAddressSpace | Generic Address Space | demonstrates the Generic Address Space feature introduced in OpenCL 2.0, which allows pointers to be declared without qualifying with a named address space |
| RangeMinimumQuery | Shared Virtual Memory pointer with offset | demonstrates passing a pointer with offset as a kernel argument using Range Minimum Query algorithm, new in OpenCL 2.0 |
| SVMAtomicsBinaryTreeInsert | SVM Fine Grain Buffer + Platform Atomics | demonstrates the Fine Grain SVM buffer with Platform atomics using a Binary Tree node insertion algorithm |
| CalcPie | C++ 11 Atomics | demonstrates atomics in OpenCL 2.0. It calculates the value of Pi using MonteCarlo analysis |
| FineGrainSVM | SVM Fine Grain Buffer + C++ 11 Atomics | demonstrates the memory model of loads and stores with new C++11 standard, which is adopted by OpenCL 2.0 (Linux APU device) |
| FineGrainSVMCAS | SVM Fine Grain Buffer + C++ 11 Atomics | demonstrates the atomic operation “CompareAndSwap” call called “atomic_compare_exchange”, introduced in OpenCL 2.0 (adopted fromC11 standards – requires Linux APU device) |
| RegionGrowingSegmentation | Device-side Enqueue | demonstrates how to use the device-side enqueue feature of OpenCL 2.0 for a Region Growing Segmentation algorithm |
| DeviceEnqueueBFS | Device-side Enqueue | demonstrates Breadth First Search implementation using the device-side enqueue feature of OpenCL 2.0. |
| ExtractPrimes | Device-side Enqueue + New Workgroup Built-in APIs | demonstrates the new workgroup builtins and device-side enqueue in a finding Prime number algorithm |
| SimpleSPIR | SPIR Consumption (Not a OpenCL 2.0 feature) | demonstrates SPIR code consumption using OpenCL APIs |
| SimpleDepthImage | Depth Image | demonstrates the depth Image APIs |
| Updated Samples | ||
| Sample | OpenCL™ 2.0 feature | Description |
| GlobalMemoryBandwidth | Shared Virtual Memory | measures the peak-bandwidth of the device buffer. For devices on OpenCL version 2.0 and higher, it additionally shows peak bandwidth for the Shared Virtual Memory (SVM) buffer |
| BinarySearch_DeviceSideEnqueue | Device-side Enqueue | enhanced to use device-side enqueue for Binary Search on an OpenCL 2.0 device. Uses iterative host side enqueue for OpenCL 1.x devices |
| BufferImageInterop | Buffer Image Interop | updated to skip checking for the BufferImageInterop extension on an OpenCL 2.0 compliant platform as it is a core feature of OpenCL 2.0 |
| BufferBandwidth | Shared Virtual Memory | now also measures the SVM buffer bandwidth on a OpenCL 2.0 compliant device |
Marty Johnson is Director of Product Engineering at AMD. His postings are his own opinions and may not represent AMD’s positions, strategies or opinions. Links to third party sites are provided for convenience and unless explicitly stated, AMD is not responsible for the contents of such linked sites and no endorsement is implied.


I can’t find any 1st hand easy to understand information about what APUs supports what level of SVM, HSA or a specific OpenCL version with what drivers and if there are any SVM opportunities (n caveats) for using dedicated CPU and GPU.
There has been sort of evolution with APUs with Kaveri being 1st GCN based APU and Carrizo with first full HSA integration.
But it is not clear if there are any differences for programming OpenCL and what the progress is with the various driver releases.
Is the http://support.amd.com/en-us/kb-articles/Pages/OpenCL2-Driver.aspx obsolete today?… what is (beyond the many game fixes) the progress with each Catalyst/Crimson driver release and why does Crimson beta for legacy GPUs not support OpenCL anymore?
There is also no information to what extend graphics drivers do make use of HSA and SVM already today. I don’t need an answer, i need a webpage that simply says what AMD HW can be used for what OpenCL and where is SVM and HSA used in drivers already. And maybe you even could “archive” old/irrelevant documentation that might lead to conFusion?