Chapter 08: Specialization Constants¶
Overview¶
Specialization constants are compile-time values baked into shaders. This chapter covers:
- Defining specialization constants in shaders
- Setting values at pipeline creation
- Use cases and performance benefits
What you'll learn:
- When specialization beats runtime branching
- Creating shader variants efficiently
- Optimizing workgroup sizes
Runtime vs Compile-Time¶
// Runtime: Branch every invocation
if (mode == 0) { /* path A */ }
else { /* path B */ }
// Compile-time: Branch removed by compiler
layout(constant_id = 0) const uint MODE = 0;
if (MODE == 0) { /* path A - only this exists in binary */ }
else { /* path B - optimized away */ }
Shader Declaration¶
#version 450
// Specialization constants with default values
layout(constant_id = 0) const uint WORKGROUP_SIZE = 256;
layout(constant_id = 1) const uint OPERATION = 0; // 0=add, 1=mul, 2=fma
layout(constant_id = 2) const float SCALE = 1.0;
layout(local_size_x_id = 0) in; // Use constant 0 for workgroup size!
layout(set = 0, binding = 0) buffer Data { float v[]; } data;
void main() {
uint idx = gl_GlobalInvocationID.x;
// Compiler optimizes away unused branches
if (OPERATION == 0) {
data.v[idx] = data.v[idx] + SCALE;
} else if (OPERATION == 1) {
data.v[idx] = data.v[idx] * SCALE;
} else {
data.v[idx] = fma(data.v[idx], SCALE, 1.0);
}
}
Note: local_size_x_id = 0 links workgroup size to specialization constant 0!
Setting Values at Pipeline Creation¶
Define the Map¶
typedef struct {
uint32_t workgroup_size;
uint32_t operation;
float scale;
} SpecConstants;
SpecConstants spec = {
.workgroup_size = 512,
.operation = 1, // multiply
.scale = 2.0f
};
VkSpecializationMapEntry entries[] = {
{
.constantID = 0,
.offset = offsetof(SpecConstants, workgroup_size),
.size = sizeof(uint32_t)
},
{
.constantID = 1,
.offset = offsetof(SpecConstants, operation),
.size = sizeof(uint32_t)
},
{
.constantID = 2,
.offset = offsetof(SpecConstants, scale),
.size = sizeof(float)
}
};
VkSpecializationInfo spec_info = {
.mapEntryCount = 3,
.pMapEntries = entries,
.dataSize = sizeof(spec),
.pData = &spec
};
Create Pipeline with Specialization¶
VkPipelineShaderStageCreateInfo stage_info = {
.sType = VK_STRUCTURE_TYPE_PIPELINE_SHADER_STAGE_CREATE_INFO,
.stage = VK_SHADER_STAGE_COMPUTE_BIT,
.module = shader_module,
.pName = "main",
.pSpecializationInfo = &spec_info // <-- Here!
};
VkComputePipelineCreateInfo pipeline_info = {
.sType = VK_STRUCTURE_TYPE_COMPUTE_PIPELINE_CREATE_INFO,
.stage = stage_info,
.layout = pipeline_layout
};
vkCreateComputePipelines(device, cache, 1, &pipeline_info, NULL, &pipeline);
Creating Shader Variants¶
Generate multiple pipelines with different constants:
typedef struct {
uint32_t workgroup_size;
uint32_t operation;
} Variant;
Variant variants[] = {
{64, 0}, // Small workgroup, add
{256, 0}, // Medium workgroup, add
{512, 0}, // Large workgroup, add
{256, 1}, // Medium workgroup, multiply
{256, 2}, // Medium workgroup, fma
};
VkPipeline pipelines[5];
for (int i = 0; i < 5; i++) {
// Update spec data
spec.workgroup_size = variants[i].workgroup_size;
spec.operation = variants[i].operation;
// Create pipeline (reuses shader module)
vkCreateComputePipelines(device, cache, 1, &pipeline_info,
NULL, &pipelines[i]);
}
Optimal Workgroup Sizes¶
Different GPUs prefer different workgroup sizes:
| GPU Vendor | Optimal Size |
|---|---|
| NVIDIA | 256, 512, 1024 |
| AMD | 64, 256 |
| Intel | 16, 32 |
| Apple | 256, 1024 |
Use specialization to create vendor-optimized pipelines:
uint32_t optimal_size;
if (strstr(props.deviceName, "NVIDIA")) {
optimal_size = 256;
} else if (strstr(props.deviceName, "AMD")) {
optimal_size = 64;
} else {
optimal_size = 128; // Safe default
}
spec.workgroup_size = optimal_size;
Running the Example¶
Output from an Apple M3 Pro, abridged:
=== Device Limits ===
Max Workgroup Size X: 1024
Max Workgroup Invocations: 1024
=== Creating Specialized Pipelines ===
Created: Add, WG=64 (2.50 ms)
Created: Add, WG=256 (0.75 ms)
Created: Add, WG=512 (0.81 ms)
Created: Multiply, WG=256 (0.82 ms)
Created: FMA, WG=256 (0.64 ms)
Created: FMA, WG=256, UF=4 (0.69 ms)
=== Benchmarking Specialized Pipelines ===
Add, WG=64:
Time: 0.261 ms/iter, Throughput: 11.23 GB/s
Verify: result[0] = 262144.0 (expected 262144.0) [ok]
Add, WG=256:
Time: 0.181 ms/iter, Throughput: 16.16 GB/s
Verify: result[0] = 262144.0 (expected 262144.0) [ok]
Add, WG=512:
Time: 0.167 ms/iter, Throughput: 17.55 GB/s
Verify: result[0] = 262144.0 (expected 262144.0) [ok]
...
=== Operation Results Comparison ===
Input: a[5] = 5.0, b[5] = 262139.0, scale = 2.0
Add, WG=64: result[5] = 262144.0
Add, WG=256: result[5] = 262144.0
Add, WG=512: result[5] = 262144.0
Multiply, WG=256: result[5] = 1310695.0
FMA, WG=256: result[5] = 262149.0
FMA, WG=256, UF=4: result[5] = 262149.0
Chapter 08 completed!
Each variant is the same SPIR-V module specialized differently at pipeline creation. On this GPU the workgroup size makes a modest difference (WG=64 is the slowest); the ranking is hardware-dependent, so measure on your target.
Supported Types¶
Specialization constants support:
| GLSL Type | SPIR-V Type |
|---|---|
bool |
OpTypeBool |
int |
OpTypeInt |
uint |
OpTypeInt |
float |
OpTypeFloat |
double |
OpTypeFloat |
No Vectors/Matrices
Composite types aren't directly supported. Use separate constants:
Use Cases¶
1. Algorithm Selection¶
layout(constant_id = 0) const uint ALGORITHM = 0;
if (ALGORITHM == 0) {
// Fast approximate version
} else {
// Slow accurate version
}
2. Feature Toggles¶
layout(constant_id = 0) const bool USE_FAST_MATH = true;
layout(constant_id = 1) const bool DEBUG_OUTPUT = false;
3. Array Sizes¶
4. Loop Unrolling¶
layout(constant_id = 0) const uint ITERATIONS = 4;
// Compiler can unroll this loop
for (uint i = 0; i < ITERATIONS; i++) {
// ...
}
Specialization vs Push Constants¶
| Aspect | Specialization | Push Constants |
|---|---|---|
| When set | Pipeline creation | Command recording |
| Changeable | No (need new pipeline) | Yes (per-dispatch) |
| Performance | Best (compiled in) | Good |
| Use for | Algorithm, sizes | Parameters, matrices |
Exercises¶
-
Auto-Tuning: Create a benchmark that tests different workgroup sizes and selects the best.
-
Feature Flags: Implement debug visualization that can be compiled out in release.
-
Kernel Variants: Create convolution kernels with specialized sizes (3x3, 5x5, 7x7).
Common Errors¶
Default Value Used¶
Forgot to provide specialization info:
Constant ID Mismatch¶
Shader uses ID 0, but map entry uses ID 1:
Size Mismatch¶
Shader expects uint32, but you provide uint64:
What's Next?¶
We've optimized individual invocations. In Chapter 09, we'll learn about subgroups — cooperating groups of invocations that can communicate without barriers.