Overview
The brief for Unit 46 simulated a real commission: build a proprietary 3D
engine from nothing over two semesters, with a list of mandatory systems —
multithreading, scene graph, ECS, a modern graphics API, scripting, render
to texture, shadow mapping, deferred rendering, and one technical
extension of our choosing — under code-quality constraints of C++20, RAII,
SOLID, no raw new/delete/malloc,
and a warning-free build.
We chose Vulkan even though the module recommended OpenGL. It was harder and it cost us time early on, but it is what the industry ships, and a coursework deadline is a cheap place to pay that learning cost.
The engine compiles as a static library and the demos are executables that
include only ghost.h: user code never sees a Vulkan, GLFW or
OpenAL header. That boundary was a requirement of the brief and ended up
being the decision that organised the whole project. The proof that it
holds is the Pacman demo — a complete 3D game written entirely in Lua,
without touching a line of the engine.
Key features
- Four-pass deferred pipeline: cascaded shadows → a four-attachment G-buffer → Cook-Torrance PBR lighting (GGX, Smith, Schlick) → HDR post-processing, with a G-buffer visualiser built into the editor.
- Bindless textures: a heap of 4096 descriptors, with the texture indexed from the material through a push constant. Not a single descriptor bind per draw call.
- Cascaded shadow maps: 3 cascades at 2048×2048, linear/logarithmic split distribution (lambda 0.75) and PCF filtering.
- Multithreaded job system built on futures: loading a scene's models and textures is spread across every hardware thread.
- ECS with contiguous per-component vectors and a parent-child hierarchy whose world matrix is cached and recomputed only when something is marked dirty.
- Procedural terrain: 6-octave fBm with domain warping, ridge blend and a power curve, plus hydraulic erosion simulated with 50,000 particles, and vegetation as camera-facing billboards.
- Streaming and LOD with six-plane frustum culling and a priority multiplier derived from the size of each entity's bounding box.
- Lua 5.4 scripting through Sol2, with script hot-reload from the inspector without stopping the scene.
- OpenAL audio with 3D positioning and BranchingSound: a node graph that switches track at the next clip boundary, so the music never cuts mid-phrase.
- Full dockable editor: hierarchy with drag-and-drop reparenting, per-component inspector, an asset browser that spawns entities when you drop a file, a profiler, and scene save/load to JSON.
- Hot-swappable tonemapping — ACES, Hable/Uncharted 2, linear — with exposure and gamma.
Goals & constraints
- A 16 ms frame budget as a design target. That is what justifies the deferred pipeline, the frustum culling and the LOD streaming.
- Load a map of more than a thousand entities with hundreds of distinct textures.
- Code constraints imposed by the brief: C++20, RAII, SOLID, zero raw allocation, and a clean build under
/W4 /WX— warnings as errors. All met. - Total platform abstraction: the user never includes Vulkan, GLFW or OpenAL.
- Deliberately out of scope: physics, skeletal animation, anti-aliasing (TAA/MSAA) and SSAO.
- The real constraint was neither hardware nor scope — it was learning Vulkan from nothing, on a fixed deadline, alongside the rest of the degree.
Architecture
The engine is a static library; the demos are executables that consume it.
GHostEnginestarts the subsystems in order — window, renderer, editor, scene — and runs the main loop: poll events, delta time,beginFrame→scene.update→endFrame.- Platform abstraction.
GHRendererholds aunique_ptr<GHVulkanRenderer>behind PIMPL, so the public header contains no Vulkan type at all. The backend is selected by macro, which means adding DX12 or Metal would be a new class and no change to user code. - ECS. Entities are 32-bit IDs; components live in contiguous per-type vectors.
TransformComponentECSmaintains the parent-child tree and caches the world matrix, recomputing it only when the entity or an ancestor is dirty. - Job system. N threads blocked on a condition variable;
enqueue()wraps any callable in apackaged_taskand returns a future. It is used above all inScene::Load(), to pull models and textures in parallel. - Renderer. Four passes per frame — shadows, G-buffer, lighting, post-processing — plus the ImGui pass and the present.
- Scenes and assets. Scenes serialise to JSON (
.ghscene) in topological order, so the hierarchy rebuilds correctly. Models are deduplicated by path withshared_ptr, so two entities using the same model share one buffer in VRAM.
Technical decisions
Vulkan rather than OpenGL 4.6, which is what the brief asked for. Explicit control of GPU memory through VMA, with no driver heuristics causing frame-time spikes; thread-safe command buffers, which is what allows draws to be recorded from several threads at all; and validation layers that hand you the exact call stack, the parameter values and the specific clause of the specification you violated. The price paid was a good deal more implementation complexity.
Deferred rather than forward. We compared a forward path with 8 lights against the target complexity and it scaled as O(geometry × lights). Deferred decouples lighting cost from geometry — O(pixels × lights) — which is what matters with many dynamic lights. It is paid for in G-buffer bandwidth and in not getting transparency for free.
Forward+ was ruled out. It would have solved transparency and saved the G-buffer bandwidth, but it needs a prior light-culling pass by tiles, and that CPU complexity did not fit in the time available.
Bindless. With hundreds of distinct materials, the
alternative is a descriptor set per material and a bind per object. With
bindless, the texture is an integer inside the push constant and the
shader does texture(textures[index], uv).
An ECS with contiguous vectors rather than a classic object hierarchy. Iterating one component type walks contiguous memory instead of chasing scattered pointers. A sparse-set ECS in the style of EnTT was considered — it iterates better still — but it complicates memory management considerably when entities are destroyed.
Lua and Sol2 rather than a bespoke scripting language, which the brief also allowed. Lua is the industry's de facto standard, has a tiny runtime to embed, and Sol2 gives type-safe binding with no overhead. Writing a language would have eaten the time that went into the renderer instead.
No coroutines and no asynchronous scripting, on purpose. Every Lua update function is synchronous and has to return inside the frame budget. Anything expensive happens in C++ and reaches Lua already computed.
Premake5 rather than CMake. It generates idiomatic
.vcxproj files without CMake's generator expressions, and the
Lua model was more readable for the per-file filters we needed — third
party warning suppression, external includes.
Systems in detail
The four-pass deferred pipeline
What it does. Turns a scene with many dynamic lights into a tonemapped image, keeping the cost of lighting independent of geometric complexity.
- Shadows (CSM). Three depth-only sub-passes, one per cascade, rendering opaque geometry from the directional light's point of view. No fragment shader. A barrier at the end moves the depth array to
SHADER_READ_ONLY. - G-buffer. A single render pass with four colour attachments and one depth attachment. It writes albedo, world-space normals transformed by the TBN basis (in
R16G16B16A16_SFLOAT, to keep precision), and packed metallic/roughness/AO. The model matrix and material index travel as a push constant — 128 bytes, the maximum the specification guarantees — which is what avoids a descriptor rebind per entity. - Lighting. A single full-screen triangle with no vertex buffer. It reconstructs world position from the depth buffer (
worldPos = invViewProj * clipPos, divided by w), then iterates the lights evaluating Cook-Torrance: GGX for the normal distribution, Smith for geometry and Schlick for Fresnel. It writes HDR radiance. - Post-processing. Another full-screen triangle applying exposure, the selected tonemapping curve and gamma correction onto the swapchain image.
Why this way. The important part is step three: reconstructing position from depth instead of storing it in the G-buffer saves an entire attachment of bandwidth per frame — which is precisely the cost deferred rendering has against it. And the G-buffer's depth buffer is reused directly, with no depth pre-pass.
Procedural terrain with hydraulic erosion
What it does. Generates terrain in chunks from noise, with believable relief and vegetation, at a per-chunk resolution set by the LOD band: 64, 32, 16 or 8 vertices per side.
How it works. The height at each (x, z) runs through a chain: domain warp — displacing the sampling point with a second noise field to break repetition — then 6-octave Perlin fBm, each octave at twice the frequency and half the amplitude, then an optional ridge blend mixing in the absolute value of the noise to produce sharp crests, then a power curve controlling the split between flat and steep ground, and finally a scale into world units.
After that, if enabled, hydraulic erosion runs: 50,000 particles by default, each with a velocity vector and a sediment load, eroding where it accelerates downhill and depositing where it slows. After enough iterations you get V-shaped valleys and alluvial fans, which is what stops the terrain looking like noise.
Why this way. The first version was a single mesh and had to be thrown away: without chunks there is no way to have LOD or to cull anything. Erosion is expensive, but it is paid once at generation rather than per frame — and it is what separates terrain from a noise heightmap.
Streaming, LOD and frustum culling
What it does. Decides, several times a second, which entities are visible and at what texture quality, so a scene larger than VRAM still fits inside the frame budget.
How it works. Every 0.1 s each entity is evaluated against five LOD bands. Frustum culling comes first: the six planes are extracted from the view-projection matrix by adding and subtracting rows, normalised, and tested against the entity's AABB with the positive-vertex test — the fastest possible rejection. Anything outside is not evaluated at all.
On top of that sits a priority multiplier. A one-metre rock and a
fifty-metre building both start at LOD 0, but at twenty metres the rock can
already drop to mip 2 and the building cannot. Automatic mode computes
clamp(AABB diagonal / reference size, min, max), with the
reference at five metres.
Why this way. The initial fixed bands caused visible quality drops on large terrain chunks. The size multiplier fixed that without having to configure every entity by hand, which was the other option and does not scale.
Development process
1. Infrastructure and the first triangle
Premake5 with Conan 2, a GLFW window and a Vulkan instance. We split the tasks between the two of us and got the render flow finished to the point of drawing a triangle, then extracted the core classes out of it. The validation layers were decisive from day one: they caught badly requested extensions and queues before they turned into hangs that are impossible to debug.
2. Models, the job system and the ECS
Rendering real models with heavy polygon counts. Loading them through the job system meant huge models appeared without a long wait, and paired with the ECS we could place and transform large numbers of entities without worrying about the cost. This is where the deferred decision was taken: a forward path with 8 lights was compared against the target complexity, and O(geometry × lights) did not fit the budget. We paid bandwidth to get O(pixels × lights).
3. The editor
Building the ImGui layout into something closer to an industry tool, to move operations out of the code and into runtime: loading and unloading models, assigning textures, transforming entities, saving and loading scenes, and attaching scripts to an entity.
4. Lighting, shadows and post-processing
The hardest phase, and it was synchronisation that made it hard: chaining shadows → G-buffer → lighting → post-processing forces you to be completely clear about every image's layout transitions. The terrain was rebuilt from scratch here too — the first version was a single mesh that could not support LOD, the second moved to chunks with per-chunk resolution.
5. Editor polish and demos
A profiler panel, terrain UI, streaming controls and shadow bias. The
Pacman demo was written entirely in Lua, to prove the engine works as a
platform without touching its code. The final weeks went on quality:
turning on /W4 /WX, fixing every warning — including those
from third-party libraries, using pragmas and Premake rules — and
documenting the public API with Doxygen.
Problems & solutions
ImGui's docking branch would not run on our Vulkan 1.0 renderer
Symptom. The engine was built against Vulkan 1.0. The editor layout we wanted needed the docking version of Dear ImGui, and the two would not work together.
Diagnosis. This one was visible rather than hidden: the order in which our renderer was submitting and presenting work did not match what the docking backend expected. No profiler needed — the incompatibility was structural, in how the frame was organised.
Fix. We took the expensive decision: migrate the engine to Vulkan 1.4 and rewrite a large part of the rendering code around it. That meant many hours of restructuring with nothing visible to show for it, which is its own kind of difficult when a deadline is running.
Result. The docking editor worked. The migration also paid for itself beyond the original problem: the newer API let us delete a significant amount of code, and rebuilding that layer a second time — with what we knew by then — left us understanding Vulkan far better than the first pass ever did.
Cascaded shadows need writing layer by layer and reading all at once
Symptom. The shadow atlas is a 2D array image of three layers, one per cascade. The shadow pass has to write into a single layer per sub-pass, while the lighting pass has to sample all three at once. Vulkan will not let the same image view do both.
Diagnosis. The validation layers pointed straight at it: a
framebuffer needs a single-layer VK_IMAGE_VIEW_TYPE_2D view
as its depth attachment, and the shader's sampler needs a
VK_IMAGE_VIEW_TYPE_2D_ARRAY view carrying all three.
Fix. GHShadowMap creates five views over the
same VkImage and the same VMA allocation: three 2D views
(base array layer 0, 1 and 2, with a layer count of 1) for the three
framebuffers of the shadow pass, and one 2D_ARRAY view with all three
layers, which is what gets bound to the lighting pass descriptor. Between
them sits a barrier moving the image from
DEPTH_STENCIL_ATTACHMENT_OPTIMAL to
SHADER_READ_ONLY_OPTIMAL.
Result. CSM working with 3 cascades at 2048×2048, a far plane at 300 units and lambda 0.75 for the split distribution. Getting there took two earlier iterations: a single 4096 shadow map, with very visible aliasing near the camera, and a two-cascade version that was better up close but still aliased in the distance.
Vulkan descriptor sets are immutable once recorded, but bindless has to grow
Symptom. The bindless approach needs a descriptor set that textures are added to as they load, at any moment. But a descriptor set already bound in a recorded command buffer cannot be modified without invalidating it.
Diagnosis. A design limitation of the API rather than a bug: the conflict is between "load assets at any time" and "command buffers already recorded have to stay valid".
Fix. Reserve the bindless heap at a fixed size — 4096
slots of COMBINED_IMAGE_SAMPLER — at initialisation, and write
new textures into free slots with vkUpdateDescriptorSets.
Slots are not released during a session, only when the engine shuts down.
We accepted a hard cap on simultaneous textures in exchange for never
invalidating anything.
Result. The bindless index is stored in the
TextureComponent and reaches the shader as a push constant,
which accesses it with texture(textures[index], uv). Zero
descriptor binds per draw call, and 4096 slots have been more than enough
in every scene tested.
// GHJobSystem: enqueue any callable and get its result back through a future
template<typename F, class... Args>
auto enqueue(F&& f, Args&&... args) -> std::future<invoke_result_t<F, Args...>> {
auto task = make_shared<packaged_task<return_type()>>(
bind(forward<F>(f), forward<Args>(args)...));
future<return_type> res = task->get_future();
{ unique_lock<mutex> lock(queue_mutex_);
tasks_.emplace([task](){ (*task)(); }); }
condition_.notify_one();
return res;
}
The job system knows nothing about what it runs. Any callable with any
signature is wrapped in a packaged_task and stored as a
function<void()> in the queue; the caller keeps a future
to collect the result whenever it suits. That is what lets
Scene::Load() fire off every model and every texture in
parallel and then use each future's get() as a natural
barrier before configuring the pipeline — no hand-written locks in the
loading code, and no engine type the job system has to know about.
Results
- A four-pass frame — CSM, G-buffer, lighting, post-processing — inside the 16 ms budget on the scenes tested.
- Up to 64 dynamic lights on the GPU, and 3 shadow cascades at 2048×2048 for the directional light.
- A bindless heap of 4096 textures, with scene loading parallelised across every thread by the job system.
- 11 runnable demos, among them a playable 3D Pacman written entirely in Lua without modifying the engine, and a profiler demo with a live frame-time graph.
- A clean build under
/W4 /WXand a public API documented with Doxygen.
The honest gap is the same one as everywhere else on this site: there are no measured figures here. The engine has a working profiler panel, so this is a matter of opening it and writing three numbers down, not of building anything.
What I would do differently
- Mesh LOD is not wired up. The streaming system only applies mip bias to textures; the
mesh_lod_levelfield exists but never reaches geometry simplification at runtime. meshoptimizer is already integrated, so this is connecting work, not research. - There is no SSAO. Ambient occlusion is read from the G-buffer's material channel, so contacts and cavities depend entirely on baked AO maps.
- There is no anti-aliasing. Neither TAA nor MSAA: it renders at native resolution, and it shows on terrain edges above all.
- The shadows have artefacts on flat surfaces near the bias threshold. PCSS would improve quality at the cost of more samples.
- The scripting API is half-finished. Lua reaches Transform, Input, Light and DebugName, but not physics, sound or BranchingSound.
- And the real one: starting on Vulkan 1.0 cost us rewriting half the engine mid-project. Choosing the version after checking what each dependency needed would have saved weeks.
What I learned
- Building a multi-pass Vulkan pipeline from nothing and reasoning about explicit synchronisation: for each pass, listing which images are read and which are written, and placing the layout transition barriers that follow from it. And debugging with validation layers as a zero-tolerance policy.
- Implementing PBR properly — Cook-Torrance with GGX, Smith and Schlick — and understanding what each term does instead of copying the formula.
- Cascaded shadow maps: computing the cascade splits, the light-space matrices, per-layer image views and PCF filtering.
- Working with array images in Vulkan, and knowing when more than one view over the same
VkImageis required. - Bindless: managing a fixed-size descriptor heap and passing indices through push constants.
- Writing a job system with
packaged_task, futures and a condition variable, and actually using it to parallelise asset loading. - Designing an API with PIMPL so the user never sees the implementation, leaving the door open to another graphics backend behind a compile-time macro.
- Embedding Lua in C++ with Sol2 and exposing engine types and functions to it.
- Working under
/W4 /WXin earnest: isolating third-party warnings with pragmas and build-system rules, without lowering the global level. - Setting up a serious C++ project with Premake5 and Conan 2, with pinned dependencies.
Running it
Requirements: Visual Studio 2022, the Vulkan SDK and Conan 2.
conan install .in the root — downloads and builds the dependencies intodeps/binaries/.- Generate the solution with Premake5 from
Build.lua. - Build the solution and run any of the demos in
examples/— Sandbox is the one that shows everything.
Shaders are compiled with tools/compile_shaders.bat
(glslangValidator, GLSL 4.60 → SPIR-V).
Credits & references
Built with Guillermo Boscá Ballester. Supervised by Arnau Rosselló — ESAT, Unit 46: Advanced Rendering and Visualization, 2025/2026.
Libraries. Vulkan + VulkanMemoryAllocator, GLFW 3.4, GLM 1.0.1, Assimp 5.4.3, stb_image / stb_vorbis, meshoptimizer, Dear ImGui (docking) + ImPlot + Native File Dialog, Sol2 3.3 + Lua 5.4, OpenAL Soft, nlohmann/json 3.11.3, spdlog 1.15, Premake5, Conan 2.
Technical references.
- Akenine-Möller, T. et al. (2018). Real-Time Rendering, 4th ed. CRC Press — the basis for the Cook-Torrance implementation, chapters 9 and 14.
- Gregory, J. (2018). Game Engine Architecture, 3rd ed. CRC Press.
- Pharr, M., Jakob, W. and Humphreys, G. (2016). Physically Based Rendering, 3rd ed. Morgan Kaufmann.
- Kosarevsky, S. and Latypov, V. (2021). 3D Graphics Rendering Cookbook. Packt.
- Sellers, G., Kessenich, J. and Alexander, R. (2016). Vulkan Programming Guide. Addison-Wesley.
- Nystrom, R. (2014). Game Programming Patterns.
- Narkowicz, K. (2015). ACES tonemapping approximation for real time — the one the engine implements.
- Hable, J. The Uncharted 2 tonemapping curve.