![]() |
vkmEngine 1.0.0
A C++ game engine · vkmengine.com
|
The engine uses one singleton thread pool for data-parallel workloads. It is a shared work-queue pool, not a per-worker work-stealing design: all workers pull from the same two queues - the frame's and the background's - guarded by one mutex and two condition variables.
Singleton via ThreadPool::get(). Spawns one worker fewer than hardware_concurrency(), because the thread that calls parallelFor runs a chunk too - a worker per core would make one more thread runnable than there are cores. At least one, since the count is allowed to be unknown. Non-copyable, non-movable; the constructor and destructor are private.
One batch form: borrowed. A Batch is a count and a task called with each index; the caller owns both, the pool borrows them, and the whole batch is one queue entry that workers claim indices from one at a time - so submitting it allocates nothing, whatever its size. parallelFor keeps its batch on its stack because it blocks until the batch is done; a caller that returns before it joins keeps its batch and task on the object the tasks write into.
Waiting is per batch, never global. The pool is shared with the async asset decodes (texture, mesh, cooked) and background bakes, so a barrier on "everything in flight" would park a frame behind an unrelated file read. A batch raises its own pending counter and each index drops it as it retires, and waitForBatch sleeps until that counter, and only that counter, reaches zero. A fire-and-forget addTask carries no counter.
The waiter helps first. Before it sleeps, waitForBatch claims every index of its batch no worker has taken yet, from wherever the batch's entry stands in the queue, and runs them itself. A caller that slept instead would leave its own work waiting for a worker to come free - a core idle for as long as the pool is busy with something else.
A throw belongs to the batch, not to the thread. An index that throws still retires, and the first exception any index of the batch threw is kept on it and rethrown by waitForBatch, whichever thread ran it. An addTask has nobody waiting to rethrow to, so the worker loop catches and logs what one throws; its body is expected to report its own errors.
A retire is a lock-free decrement. Only the drop to zero takes the mutex, and only to order the notify: a waiter holds the mutex from reading a nonzero count until it is asleep, so a notify issued after taking it cannot land in between. The counter is not touched after that decrement, because its waiter may already have returned.
shutdown() joins the workers while the rest of the program is still standing. Engine::run calls it once the main loop exits: the pool is a function-local static, so its own destructor runs after the singletons an in-flight decode pushes into. It is idempotent, and parallelFor sweeps serially once the workers are gone.
Two queues, chosen by the shape of the call. Batches land in m_frameTasks and addTask in m_backgroundTasks, and a worker drains the first before it looks at the second. The reason is the counter above: somebody is blocked in waitForBatch until a batch empties and that somebody is the frame, while every addTask caller - an asset decode or a background bake - retires against nothing. The order does not starve a decode for long: a worker reaches the second queue as soon as the first holds nothing unclaimed, and only the frame fills the first - parallelFor blocks its caller until its range is done, and the one batch forked by hand (below) is joined before the pass that reads it. That batch stays out while the frame goes on, so "one batch at a time" is not the guarantee; what holds is that every batch in the queue belongs to the frame in progress.
Internals: those two std::deques of queued work - an owned task, or a batch that stays in the queue until its last index is claimed, carrying the counter its indices retire against - a m_tasksMutex, a m_tasksCV (wakes workers when tasks arrive) and a m_doneCV (signalled when a batch counter hits 0). The split is by who is waiting, not per worker: there is no local/global queue and no stealing.
parallelFor is a free function, not a method. Two overloads:
The callback is invoked per index. It may take the index (function(size_t i)) or nothing (function()); the implementation picks the form via if constexpr (is_invocable_v<Function, size_t>).
Behavior:
parallelFor is the per-system scaling lever (the framework does not parallelize systems against each other - see architecture.md). The main consumers are VisibilitySystem (culling every mesh into pre-sized per-index arrays), HierarchySystem (resolving world transforms a level of the tree at a time) and PhysicsSystem's narrowphase (colliding the broadphase's pairs a chunk at a time). For correctness, parallel work writes into pre-sized per-index or per-chunk buffers rather than sharing mutable state across the loop.
The narrowphase is the one consumer that merges rather than writing in place: a pair yields any number of manifolds, and several pairs share a body. Each chunk writes only its own buffers, kept on the system and cleared each tick, and a serial merge after the join reads them in chunk order, so what reaches the solver, the bodies and the event bus is in pair order whichever thread ran which chunk (physics.md).
One consumer forks and joins by hand rather than through parallelFor: GLShadowData::cullCasters queues one batch, an index per cascade, spot and cube face, and returns, and GLShadowData::finishCull is the waitForBatch, called by the backend just before the first pass. Everything the backend builds between the two - UBOs, the drawable partition, the skin palette, the opaque batch - reads none of the batches, so it runs beside the cull instead of after it. It is the same fork-join, with the join placed where the result is first needed; the batch and its task live on the object the tasks write into, and build and the destructor both join first so no task ever outlives what it writes.