Mesh Systems and Vertex Packing
This document provides a complete, replicable guide to Pyrite’s mesh architecture, including data structures, vertex packing, the greedy meshing algorithm, and ambient occlusion (AO) calculation.
Overview
Pyrite renders millions of voxels using a multi-stage pipeline where lighting is generated first, then meshed/packed, then finally interpreted by shaders:
Voxel data availability: chunk voxels + (optional) chunk lightmap.
Lighting generation (elsewhere): sunlight + blocklight are computed by the lighting system and stored in the per-chunk lightmap.
Greedy Meshing: CPU-side algorithm groups adjacent coplanar faces into large rectangular polygons.
Vertex Packing: Vertex attributes compressed into 32-bit integers to minimize GPU memory.
AO + per-vertex light smoothing: - AO is computed from neighboring blocks (baked into the vertex). - Sunlight/blocklight values from the chunk lightmap are sampled and averaged into a per-vertex light_data value.
GPU Upload: Main thread creates VAO/VBO objects from packed data.
Rendering: Draw calls per mesh (opaque + transparent passes).
Mesh Classes Hierarchy
Pyrite implements 6 distinct mesh classes to handle rendering. Each extends a foundational base class and handles specific GPU data formatting.
1. BaseMesh (Core Abstract Class)
Handles standardizing the Vertex Array Objects (VAOs) and Vertex Buffer Objects (VBOs) through the ModernGL pipeline.
self.vbo_format: str = ''
self.attrs: Tuple[str, ...] = ()
Format Mapping: Every mesh requires a format string telling ModernGL how to read the bytes (e.g.,
'3f 3f'for two vec3s) and the exact variable names used in the GLSL shader (e.g.,'in_position', 'in_color').
vbo = self.ctx.buffer(vertex_data)
vao = self.ctx.vertex_array(self.program, [(vbo, self.vbo_format, *self.attrs)])
GPU Uploading:
get_vao()takes the raw Numpy array generated by subclasses, pushes it to the GPU as a VBO, and binds it to a VAO mapped to the specific shader program.
2. ChunkMesh (Per-Chunk Geometry)
Manages the dense, greedy-meshed geometry for the voxel world.
self.vbo_format: str = '1u4 1u4'
self.attrs: Tuple[str, ...] = ('packed_data', 'light_data')
Data Payload:
1u4means “one unsigned 4-byte integer” (uint32). The chunk mesh passes exactly two 32-bit integers per vertex to the shader: the bit-packed position/ID payload, and the packed sunlight/blocklight payload.
pool.rotate(-best_i)
self.vbo, self.vao = pool.popleft()
self.vbo.write(self.vertex_data)
VBO Pooling (Crucial Optimization): Instead of destroying and allocating new GPU memory every time a player moves, Pyrite searches the
world.vbo_poolfor an unused buffer that is large enough to fit the new chunk data. It re-writes the data over the old memory, completely eliminating VRAM leaks!
reserve_size = 2 ** math.ceil(math.log2(byte_size)) if byte_size > 0 else 0
self.vbo = self.ctx.buffer(reserve=reserve_size)
Power-of-2 Allocation: If a new VBO must be created, Pyrite mathematically rounds the requested byte size up to the nearest power of 2 (e.g., 64KB, 128KB). This standardizes buffer sizes, making them highly reusable for future chunks.
self.vao.render(vertices=self.water_count, first=self.opaque_count)
Multi-Pass Rendering: Because opaque and transparent faces are packed into the same VBO buffer,
render_water()tells the GPU to start drawing water faces exactly where the opaque vertices end.
3. CloudMesh (Procedural Clouds)
Generates the scrolling 3D cloud layer.
self.vbo_format: str = '3u2'
16-bit Precision:
3u2means “three unsigned 2-byte integers” (uint16). Clouds only need basic X, Y, Z coordinates. Using 16-bit integers instead of 32-bit floats cuts cloud VRAM usage perfectly in half.
while x + x_count < width and cloud_data[idx] and idx not in visited:
x_count += 1
2D Greedy Meshing: The clouds are generated as a flat 2D density map via Simplex Noise. The mesh builder scans along the X-axis, grouping adjacent cloud blocks into a single strip. It then extends that strip along the Z-axis to form a massive rectangular polygon, drastically reducing cloud draw calls.
4. CubeMesh (Voxel Marker)
Constructs a basic 1x1x1 cube used to highlight targeted blocks.
data: List[float] = [vertices[ind] for triangle in indices for ind in triangle]
Flattening: ModernGL requires flat 1D arrays. This list comprehension takes a hardcoded list of unique corner vertices and a list of triangle index pairs, and flattens them into a single contiguous array.
combined_data: NDArray[np.float16] = np.hstack([tex_coord_data, vertex_data])
Interleaving Data:
np.hstackhorizontally stacks the UV coordinates array and the 3D position array together so they interleave perfectly for the GPU.
5. ItemMesh (Dropped Items)
Builds standard cubes for dropped items with explicitly mapped face IDs.
self.vbo_format: str = '3f 2f 1f'
self.attrs: Tuple[str, ...] = ('in_position', 'in_tex_coord', 'in_face_id')
Standard Precision: Dropped items use standard 32-bit floats for position (
3f), texture UVs (2f), and the face ID (1f).
# format: position(3), uv(2), face_id(1)
# Top: 0, 1, 1, 0, 1, 0...
Explicit Mapping: The vertex data is entirely hardcoded. By explicitly tagging each vertex with a
face_id(0 to 5), the item shader can dynamically calculate the correct texture layer and shading multipliers without needing greedy meshing algorithms.
6. ObjMesh (Wavefront Models)
Parses standard .obj files to render custom 3D tools in-hand.
if line.startswith('newmtl'):
current_material = line.split()[1]
elif line.startswith('Kd'):
materials[current_material]['Kd'] = [float(x) for x in line.split()[1:]]
Material Parsing: Parses the companion
.mtlfile.Kdrepresents the diffuse color vector. Pyrite extracts these RGB values to color the 3D model vertices naturally without requiring a UV texture map.
for face in (face_vertices[0], face_vertices[i], face_vertices[i + 1]):
N-Gon Triangulation: Standard OBJ files often contain Quads (4 vertices) or N-gons. Since OpenGL only draws triangles, this loop iterates through the polygon and connects the first vertex to every subsequent pair of vertices (Triangle Fan), safely triangulating complex faces on the fly.
center_x = (np.max(x_coords) + np.min(x_coords)) / 2.0
vd_array[0::11] -= center_x
Auto-Centering: If a 3D model was exported from Blender with its origin point far away from the geometry, rotating it in-game would cause it to swing in massive, chaotic orbits. Pyrite finds the mathematical center of the bounding box and subtracts it from every vertex (using NumPy slicing
[0::11]to hit just the X coordinates), perfectly centering the model at(0, 0, 0).
Vertex Packing: 32-Bit Format
To minimize GPU bandwidth and memory, each vertex attribute is bit-packed into a single 32-bit unsigned integer.
Packed Vertex Layout:
Bits 31-26 (6 bits): X coordinate (0-47)
Bits 25-20 (6 bits): Y coordinate (0-47)
Bits 19-14 (6 bits): Z coordinate (0-47)
Bits 13-6 (8 bits): Voxel ID (0-255)
Bits 5-3 (3 bits): Face ID (0-5, one of 6 faces)
Bits 2-1 (2 bits): AO ID (0-3, ambient occlusion level)
Bit 0 (1 bit): Flip ID (0 or 1, diagonal flip flag)
Total: 32 bits = 4 bytes per vertex (vs. 16 bytes for traditional (x, y, z, id, face, ao, flip, light))
Packing Formula Breakdown:
(x & 0x3F) << 26
Coordinate Masking: The
& 0x3Fis a bitwise AND mask (63 in decimal) that strictly bounds the value to 6 bits (0-63). The<< 26shifts these 6 bits all the way to the far left of the 32-bit integer.
| (voxel_id & 0xFF) << 6
Attribute Merging: The
|(bitwise OR) securely merges this new value into the existingpacked_datainteger without overwriting the previously shifted coordinate bits.0xFFmasks exactly 8 bits for the voxel ID (0-255).
Unpacking (in Vertex Shader):
void unpack(uint packed_data) {
x = int((packed_data >> 26) & 0x3F);
y = int((packed_data >> 20) & 0x3F);
z = int((packed_data >> 14) & 0x3F);
voxel_id = int((packed_data >> 6) & 0xFF);
face_id = int((packed_data >> 3) & 0x7);
ao_id = int((packed_data >> 1) & 0x3);
flip_id = int(packed_data & 0x1);
}
Light Data (Separate uint32):
Bits 7-4 (4 bits): Sunlight (0-15)
Bits 3-0 (4 bits): Blocklight (0-15)
Greedy Meshing Algorithm
Greedy meshing reduces face count by grouping coplanar, identical-ID faces into rectangles. Executed on CPU; results are packed and uploaded to GPU.
High-Level Steps:
For each of 3 orthogonal planes (XY, XZ, YZ):
Iterate through all slices perpendicular to that plane
Build 2D mask of solid vs. transparent voxels
For each solid voxel with exposed face:
Calculate AO and light for all 4 corners
Find greedy horizontal rectangle width
Find greedy vertical rectangle height
Emit quad vertices
Mark processed faces to avoid double-processing
Separate opaque and water faces into independent buffers
Detailed Algorithm: X-Plane Scanning
Processing YZ-plane slices (X varying):
for x_slice in range(CHUNK_SIZE):
mask = np.zeros((CHUNK_SIZE, CHUNK_SIZE), dtype=np.bool_)
Mask Initialization: The mesher iterates through flat 2D slices, creating a boolean tracking array to isolate exclusively solid and exposed voxel faces.
while z + width < CHUNK_SIZE and mask[y, z + width]:
width += 1
Greedy Expansion: The algorithm continuously sweeps rightwards along the active axis, incrementing the geometric rectangle width perfectly flush as long as adjacent faces share identical properties.
Y and Z Plane Scanning work similarly, iterating through XZ and XY slices respectively.
Performance Note: This is a hot loop executed once per chunk load. Implemented in Numba with @njit(cache=True, nogil=True) for 500x+ speedup.
Vertex Light Smoothing
Light values are interpolated to vertices for smooth shading. Each vertex is shared by up to 8 blocks; we sample light from the 4 (on a plane) or 8 blocks surrounding that vertex.
For X-Plane Face (perpendicular normal = +X):
Four corners of the quad correspond to YZ positions. For each corner, sample from 4 blocks:
l1 = get_light(x, y, z) # Lower-left
l2 = get_light(x, y+1, z) # Upper-left
Corner Sampling: A vertex is shared by up to 4 faces on a 2D plane. We retrieve the raw lighting integer for all 4 surrounding blocks.
avg_sun = ((l1 >> 4) + (l2 >> 4) + (l3 >> 4) + (l4 >> 4)) / 4
Extraction and Averaging: We bitshift right by 4 (
>> 4) to extract just the sunlight values (which occupy the upper 4 bits of the light byte) from the 4 surrounding blocks, then sum and divide by 4 to get a smooth average.
return (int(avg_sun) << 4) | int(avg_block)
Repacking: We shift the averaged sunlight back up by 4 bits (
<< 4) and use a bitwise OR|to securely pack it together with the averaged blocklight into a single, smoothed byte.
Ambient Occlusion (AO) Calculation
AO darkens corners where multiple solid blocks converge, simulating soft shadows.
Corner Occlusion Breakdown:
For each of the 4 corners of a quad, check 2x2 adjacent blocks:
if not is_transparent(voxel_at(corner_y-1, corner_z-1)):
ao_count += 1
Neighbor Checking: For a specific vertex corner, we check the 2x2 grid of neighboring blocks. If a neighbor is solid, it physically blocks light, so we increment the
ao_countpenalty.
return min(ao_count, 3)
Overflow Prevention: The vertex format only has 2 bits allocated for AO, meaning it can only store values 0, 1, 2, or 3. Even though 4 neighbors could theoretically be solid, we use
min()to clamp the count. This physically prevents the value from exceeding 3, guaranteeing we don’t accidentally overwrite neighboring bits during the bit-packing phase!
Transparency Check:
Transparent blocks (AIR, WATER, GLASS, LEAVES) do not cast AO shadows:
def is_transparent(voxel_id):
return voxel_id in [AIR, WATER, GLASS, LEAVES]
GPU Application (in Vertex Shader):
const float ao_values[4] = float[4](0.1, 0.25, 0.5, 1.0);
// Unpack ao_id (2 bits)
int ao_id = int((packed_data >> 1) & 0x3);
// Apply to shading
shading = base_light * ao_values[ao_id];
Flip Detection (Diagonal Flip for Lighting)
When lighting is uneven across a quad, flipping the diagonal can improve visual appearance. This is determined by comparing lighting sums across the two diagonals.
Algorithm Breakdown:
diag1_brightness = extract_light(l0) + extract_light(l2) + (ao0 + ao2)
Diagonal Summation: To find the smoothest lighting gradient across a square face, we calculate the combined brightness (Sunlight + Blocklight + AO) of the two opposite corners (Diagonal 1) and compare it against the remaining two opposite corners (Diagonal 2).
return diag1_brightness > diag2_brightness
Triangulation Flipping: We compare the two diagonals. If Diagonal 1 is brighter, we return
True(1), signaling the mesh builder to “flip” the internal edge connecting the two triangles that make up the quad. This mathematically ensures the bright corners share a triangle face, creating a beautifully smooth visual shadow gradient rather than a harsh graphical line.
GPU Application:
During rendering, the vertex shader uses flip_id to adjust vertex positions or UV coordinates accordingly.
Water Faces Handling
Water is rendered separately to allow transparency blending without depth-test complications.
Algorithm:
During greedy meshing, water faces (voxel_id == WATER) are marked separately
Opaque faces emitted first, water faces appended to same buffer
Render call splits: draw opaque faces first (full depth test), then draw water faces (transparency blending enabled)
Separate render call (in Shader Program):
ctx.enable(moderngl.BLEND)
ctx.blend_func = (moderngl.SRC_ALPHA, moderngl.ONE_MINUS_SRC_ALPHA)
vao.render(mode=moderngl.TRIANGLES, vertices=water_count, first=opaque_count)
Mesh Building Pipeline (CPU to GPU)
Sequential Process:
Chunk Load (Background Thread):
Generate or fetch voxel data from database
Place in
load_queue
Mesh Build (Background Thread via ThreadPoolExecutor):
Pop chunk from
load_queueRun greedy meshing:
build_chunk_mesh(chunk_voxels, chunk_lightmap, format_size, chunk_pos, world_voxels, world_lightmaps, chunk_positions)Output:
vertex_data(flat uint32 array),light_data(flat uint32 array)Place in
build_queue
GPU Upload (Main Thread):
Pop from
mesh_queue(result of lighting stitching inbuild_queue)Create VAO/VBO:
ctx.vertex_array(program, vbo, vao)Store in
chunk.meshobjectIf VBO pool available, reuse; else allocate new
Rendering (Main Thread, per frame):
Frustum cull active chunks
Occlusion query invisible chunks
Bind shader, draw visible chunk meshes
VBO Pool (Memory Recycling):
if self.vbo_pool:
vbo = self.vbo_pool.pop()
vbo.write(data)
VBO Recycling: The engine pops stale buffers off the deque queue and overwrites the exact memory addresses in VRAM instantaneously without expensive destructions natively.
Data Flow Example
Raw Chunk Voxels (1D array, 110,592 elements)
↓
[Greedy Meshing: CPU]
↓
Packed Vertex Data (e.g., 10,000 vertices for a grass chunk)
↓
[Lighting Stitching: CPU]
↓
Light Data (10,000 light values)
↓
[GPU Upload: Main Thread]
↓
VBO/VAO allocated on GPU
↓
[Rendering: per frame]
↓
Vertices unpacked in Vertex Shader → Position + Attributes
↓
Fragment Shader colors pixels
Next Steps
Now that the vertex data is packed and ready on the GPU, dive into the overall Rendering Pipeline pipeline to see how frustum culling and hardware occlusion determine what actually gets drawn.