All work

Ghost Engine

A complete 3D engine written from scratch in C++20 and Vulkan: a four-pass deferred pipeline with cascaded shadow maps, bindless textures, an ECS running on a multithreaded job system, procedural terrain with hydraulic erosion, and a dockable editor with Lua scripting.

Year
2025–2026
Duration
~7 months, part time — two academic semesters
Team
2 programmers — Guillermo Boscá Ballester and myself
My areas
Deferred pipeline and the four-attachment G-buffer, Cook-Torrance PBR lighting, cascaded shadow maps, HDR post-processing and tonemapping, vegetation billboards, the job system, procedural terrain with erosion, and Lua scripting through Sol2
Stack
C++20, Vulkan (VMA 3.1), GLFW 3.4, GLM 1.0.1, Assimp 5.4.3, meshoptimizer, Sol2 3.3 + Lua 5.4, OpenAL Soft, Dear ImGui (docking) + ImPlot, nlohmann/json, spdlog, Premake5, Conan 2
Platform
Windows x64, Visual Studio 2022
Context
Unit 46, Advanced Rendering and Visualization — final year of the HND at ESAT, 2025/2026
Status
Finished

Overview

The brief for Unit 46 simulated a real commission: build a proprietary 3D engine from nothing over two semesters, with a list of mandatory systems — multithreading, scene graph, ECS, a modern graphics API, scripting, render to texture, shadow mapping, deferred rendering, and one technical extension of our choosing — under code-quality constraints of C++20, RAII, SOLID, no raw new/delete/malloc, and a warning-free build.

We chose Vulkan even though the module recommended OpenGL. It was harder and it cost us time early on, but it is what the industry ships, and a coursework deadline is a cheap place to pay that learning cost.

The engine compiles as a static library and the demos are executables that include only ghost.h: user code never sees a Vulkan, GLFW or OpenAL header. That boundary was a requirement of the brief and ended up being the decision that organised the whole project. The proof that it holds is the Pacman demo — a complete 3D game written entirely in Lua, without touching a line of the engine.

Key features

Goals & constraints

Architecture

The engine is a static library; the demos are executables that consume it.

Diagram of the engine frame loop, from startup through the main loop to shutdown
The engine lifecycle. Inside the main loop, asset loading branches off the scene update into the job system the first time it is needed, instead of blocking — which is what keeps the editor responsive while a large map streams in.

Technical decisions

Vulkan rather than OpenGL 4.6, which is what the brief asked for. Explicit control of GPU memory through VMA, with no driver heuristics causing frame-time spikes; thread-safe command buffers, which is what allows draws to be recorded from several threads at all; and validation layers that hand you the exact call stack, the parameter values and the specific clause of the specification you violated. The price paid was a good deal more implementation complexity.

Deferred rather than forward. We compared a forward path with 8 lights against the target complexity and it scaled as O(geometry × lights). Deferred decouples lighting cost from geometry — O(pixels × lights) — which is what matters with many dynamic lights. It is paid for in G-buffer bandwidth and in not getting transparency for free.

Forward+ was ruled out. It would have solved transparency and saved the G-buffer bandwidth, but it needs a prior light-culling pass by tiles, and that CPU complexity did not fit in the time available.

Bindless. With hundreds of distinct materials, the alternative is a descriptor set per material and a bind per object. With bindless, the texture is an integer inside the push constant and the shader does texture(textures[index], uv).

An ECS with contiguous vectors rather than a classic object hierarchy. Iterating one component type walks contiguous memory instead of chasing scattered pointers. A sparse-set ECS in the style of EnTT was considered — it iterates better still — but it complicates memory management considerably when entities are destroyed.

Lua and Sol2 rather than a bespoke scripting language, which the brief also allowed. Lua is the industry's de facto standard, has a tiny runtime to embed, and Sol2 gives type-safe binding with no overhead. Writing a language would have eaten the time that went into the renderer instead.

No coroutines and no asynchronous scripting, on purpose. Every Lua update function is synchronous and has to return inside the frame budget. Anything expensive happens in C++ and reaches Lua already computed.

Premake5 rather than CMake. It generates idiomatic .vcxproj files without CMake's generator expressions, and the Lua model was more readable for the per-file filters we needed — third party warning suppression, external includes.

Editor in wireframe mode with the G-buffer visualiser showing albedo, normal, material and depth targets
The deferred pipeline in one frame: the G-buffer visualiser on the left shows albedo, normals, material and depth as separate targets, with the terrain in wireframe in the viewport. The terrain's height bands — water, sand, grass and snow — are editable from the inspector.

Systems in detail

The four-pass deferred pipeline

What it does. Turns a scene with many dynamic lights into a tonemapped image, keeping the cost of lighting independent of geometric complexity.

  1. Shadows (CSM). Three depth-only sub-passes, one per cascade, rendering opaque geometry from the directional light's point of view. No fragment shader. A barrier at the end moves the depth array to SHADER_READ_ONLY.
  2. G-buffer. A single render pass with four colour attachments and one depth attachment. It writes albedo, world-space normals transformed by the TBN basis (in R16G16B16A16_SFLOAT, to keep precision), and packed metallic/roughness/AO. The model matrix and material index travel as a push constant — 128 bytes, the maximum the specification guarantees — which is what avoids a descriptor rebind per entity.
  3. Lighting. A single full-screen triangle with no vertex buffer. It reconstructs world position from the depth buffer (worldPos = invViewProj * clipPos, divided by w), then iterates the lights evaluating Cook-Torrance: GGX for the normal distribution, Smith for geometry and Schlick for Fresnel. It writes HDR radiance.
  4. Post-processing. Another full-screen triangle applying exposure, the selected tonemapping curve and gamma correction onto the swapchain image.

Why this way. The important part is step three: reconstructing position from depth instead of storing it in the G-buffer saves an entire attachment of bandwidth per frame — which is precisely the cost deferred rendering has against it. And the G-buffer's depth buffer is reused directly, with no depth pre-pass.

A high-detail scanned stone lion statue lit by two point lights in the Sandbox demo, one of them tinted violet
The lighting pass doing its job. A dense scanned mesh in the Sandbox demo, lit by two point lights — the violet one raking in from the left — with nothing but the Cook-Torrance response to carry the stone: the roughness breaking up the highlights across the mane, and the falloff picking out the carved detail.

Procedural terrain with hydraulic erosion

What it does. Generates terrain in chunks from noise, with believable relief and vegetation, at a per-chunk resolution set by the LOD band: 64, 32, 16 or 8 vertices per side.

How it works. The height at each (x, z) runs through a chain: domain warp — displacing the sampling point with a second noise field to break repetition — then 6-octave Perlin fBm, each octave at twice the frequency and half the amplitude, then an optional ridge blend mixing in the absolute value of the noise to produce sharp crests, then a power curve controlling the split between flat and steep ground, and finally a scale into world units.

After that, if enabled, hydraulic erosion runs: 50,000 particles by default, each with a velocity vector and a sediment load, eroding where it accelerates downhill and depositing where it slows. After enough iterations you get V-shaped valleys and alluvial fans, which is what stops the terrain looking like noise.

Why this way. The first version was a single mesh and had to be thrown away: without chunks there is no way to have LOD or to cull anything. Erosion is expensive, but it is paid once at generation rather than per frame — and it is what separates terrain from a noise heightmap.

Streaming, LOD and frustum culling

What it does. Decides, several times a second, which entities are visible and at what texture quality, so a scene larger than VRAM still fits inside the frame budget.

How it works. Every 0.1 s each entity is evaluated against five LOD bands. Frustum culling comes first: the six planes are extracted from the view-projection matrix by adding and subtracting rows, normalised, and tested against the entity's AABB with the positive-vertex test — the fastest possible rejection. Anything outside is not evaluated at all.

On top of that sits a priority multiplier. A one-metre rock and a fifty-metre building both start at LOD 0, but at twenty metres the rock can already drop to mip 2 and the building cannot. Automatic mode computes clamp(AABB diagonal / reference size, min, max), with the reference at five metres.

Why this way. The initial fixed bands caused visible quality drops on large terrain chunks. The size multiplier fixed that without having to configure every entity by hand, which was the other option and does not scale.

Development process

1. Infrastructure and the first triangle

Premake5 with Conan 2, a GLFW window and a Vulkan instance. We split the tasks between the two of us and got the render flow finished to the point of drawing a triangle, then extracted the core classes out of it. The validation layers were decisive from day one: they caught badly requested extensions and queues before they turned into hangs that are impossible to debug.

2. Models, the job system and the ECS

Rendering real models with heavy polygon counts. Loading them through the job system meant huge models appeared without a long wait, and paired with the ECS we could place and transform large numbers of entities without worrying about the cost. This is where the deferred decision was taken: a forward path with 8 lights was compared against the target complexity, and O(geometry × lights) did not fit the budget. We paid bandwidth to get O(pixels × lights).

3. The editor

Building the ImGui layout into something closer to an industry tool, to move operations out of the code and into runtime: loading and unloading models, assigning textures, transforming entities, saving and loading scenes, and attaching scripts to an entity.

4. Lighting, shadows and post-processing

The hardest phase, and it was synchronisation that made it hard: chaining shadows → G-buffer → lighting → post-processing forces you to be completely clear about every image's layout transitions. The terrain was rebuilt from scratch here too — the first version was a single mesh that could not support LOD, the second moved to chunks with per-chunk resolution.

5. Editor polish and demos

A profiler panel, terrain UI, streaming controls and shadow bias. The Pacman demo was written entirely in Lua, to prove the engine works as a platform without touching its code. The final weeks went on quality: turning on /W4 /WX, fixing every warning — including those from third-party libraries, using pragmas and Premake rules — and documenting the public API with Doxygen.

Problems & solutions

ImGui's docking branch would not run on our Vulkan 1.0 renderer

Symptom. The engine was built against Vulkan 1.0. The editor layout we wanted needed the docking version of Dear ImGui, and the two would not work together.

Diagnosis. This one was visible rather than hidden: the order in which our renderer was submitting and presenting work did not match what the docking backend expected. No profiler needed — the incompatibility was structural, in how the frame was organised.

Fix. We took the expensive decision: migrate the engine to Vulkan 1.4 and rewrite a large part of the rendering code around it. That meant many hours of restructuring with nothing visible to show for it, which is its own kind of difficult when a deadline is running.

Result. The docking editor worked. The migration also paid for itself beyond the original problem: the newer API let us delete a significant amount of code, and rebuilding that layer a second time — with what we knew by then — left us understanding Vulkan far better than the first pass ever did.

Cascaded shadows need writing layer by layer and reading all at once

Symptom. The shadow atlas is a 2D array image of three layers, one per cascade. The shadow pass has to write into a single layer per sub-pass, while the lighting pass has to sample all three at once. Vulkan will not let the same image view do both.

Diagnosis. The validation layers pointed straight at it: a framebuffer needs a single-layer VK_IMAGE_VIEW_TYPE_2D view as its depth attachment, and the shader's sampler needs a VK_IMAGE_VIEW_TYPE_2D_ARRAY view carrying all three.

Fix. GHShadowMap creates five views over the same VkImage and the same VMA allocation: three 2D views (base array layer 0, 1 and 2, with a layer count of 1) for the three framebuffers of the shadow pass, and one 2D_ARRAY view with all three layers, which is what gets bound to the lighting pass descriptor. Between them sits a barrier moving the image from DEPTH_STENCIL_ATTACHMENT_OPTIMAL to SHADER_READ_ONLY_OPTIMAL.

Result. CSM working with 3 cascades at 2048×2048, a far plane at 300 units and lambda 0.75 for the split distribution. Getting there took two earlier iterations: a single 4096 shadow map, with very visible aliasing near the camera, and a two-cascade version that was better up close but still aliased in the distance.

Vulkan descriptor sets are immutable once recorded, but bindless has to grow

Symptom. The bindless approach needs a descriptor set that textures are added to as they load, at any moment. But a descriptor set already bound in a recorded command buffer cannot be modified without invalidating it.

Diagnosis. A design limitation of the API rather than a bug: the conflict is between "load assets at any time" and "command buffers already recorded have to stay valid".

Fix. Reserve the bindless heap at a fixed size — 4096 slots of COMBINED_IMAGE_SAMPLER — at initialisation, and write new textures into free slots with vkUpdateDescriptorSets. Slots are not released during a session, only when the engine shuts down. We accepted a hard cap on simultaneous textures in exchange for never invalidating anything.

Result. The bindless index is stored in the TextureComponent and reaches the shader as a push constant, which accesses it with texture(textures[index], uv). Zero descriptor binds per draw call, and 4096 slots have been more than enough in every scene tested.

// GHJobSystem: enqueue any callable and get its result back through a future
template<typename F, class... Args>
auto enqueue(F&& f, Args&&... args) -> std::future<invoke_result_t<F, Args...>> {
    auto task = make_shared<packaged_task<return_type()>>(
        bind(forward<F>(f), forward<Args>(args)...));
    future<return_type> res = task->get_future();
    { unique_lock<mutex> lock(queue_mutex_);
      tasks_.emplace([task](){ (*task)(); }); }
    condition_.notify_one();
    return res;
}

The job system knows nothing about what it runs. Any callable with any signature is wrapped in a packaged_task and stored as a function<void()> in the queue; the caller keeps a future to collect the result whenever it suits. That is what lets Scene::Load() fire off every model and every texture in parallel and then use each future's get() as a natural barrier before configuring the pipeline — no hand-written locks in the loading code, and no engine type the job system has to know about.

Results

The honest gap is the same one as everywhere else on this site: there are no measured figures here. The engine has a working profiler panel, so this is a matter of opening it and writing three numbers down, not of building anything.

Lit terrain scene with a red point light, inspector showing billboard and light components, and the frame profiler
A point light over the terrain. The inspector shows the billboard and light components on the selected entity — including its bindless texture index — and the profiler panel on the left tracks current, average, minimum and maximum frame time.

What I would do differently

What I learned

Running it

Requirements: Visual Studio 2022, the Vulkan SDK and Conan 2.

  1. conan install . in the root — downloads and builds the dependencies into deps/binaries/.
  2. Generate the solution with Premake5 from Build.lua.
  3. Build the solution and run any of the demos in examples/ — Sandbox is the one that shows everything.

Shaders are compiled with tools/compile_shaders.bat (glslangValidator, GLSL 4.60 → SPIR-V).

Credits & references

Built with Guillermo Boscá Ballester. Supervised by Arnau Rosselló — ESAT, Unit 46: Advanced Rendering and Visualization, 2025/2026.

Libraries. Vulkan + VulkanMemoryAllocator, GLFW 3.4, GLM 1.0.1, Assimp 5.4.3, stb_image / stb_vorbis, meshoptimizer, Dear ImGui (docking) + ImPlot + Native File Dialog, Sol2 3.3 + Lua 5.4, OpenAL Soft, nlohmann/json 3.11.3, spdlog 1.15, Premake5, Conan 2.

Technical references.