{ "root_post_id": "1858715434436227544", "posts": [ { "post_id": "1858709400791261630", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "Working towards a new bring up of front-buffering on NV via VK. My last implementation had been tuned on AMD and didn't work on NV any more. NV does seem to accept a 1-deep swap with IMMEDIATE presentation at least on latest drivers, so that is a good start.", "timestamp": "2024-11-19 03:10:28", "media_urls": [], "reply_to_id": null, "quote_of_id": null, "metrics": { "reply_count": 1, "repost_count": 0, "like_count": 2, "view_count": 686 } }, { "post_id": "1858710122727391666", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "Since my image resources are static after init time, only swap images change when the driver kills the swap chain (which once upon a time seemed to happen on ALT+TAB maybe). So it's one descriptor set always bound.", "timestamp": "2024-11-19 03:13:21", "media_urls": [], "reply_to_id": "1858709400791261630", "quote_of_id": null, "metrics": { "reply_count": 1, "repost_count": 0, "like_count": 0, "view_count": 86 } }, { "post_id": "1858710719018979651", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "Stopped using VK_DESCRIPTOR_SET_LAYOUT_CREATE_UPDATE_AFTER_BIND_POOL_BIT_EXT (god awful naming length people), because I think it could be a perf hit on NV due to extra indirection. NV driver is free to bake down the single set now before the command buffer is sent over.", "timestamp": "2024-11-19 03:15:43", "media_urls": [], "reply_to_id": "1858710122727391666", "quote_of_id": null, "metrics": { "reply_count": 1, "repost_count": 0, "like_count": 0, "view_count": 105 } }, { "post_id": "1858712158004994552", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "At this stage, to the point of having everything loaded and swap chain created, it's 2700 lines of engine code. That includes embedding headers (no external includes), so it's a one file compile. People claim rolling your own engine is hard? Not really if you keep it focused.", "timestamp": "2024-11-19 03:21:26", "media_urls": [], "reply_to_id": "1858710719018979651", "quote_of_id": null, "metrics": { "reply_count": 1, "repost_count": 0, "like_count": 2, "view_count": 101 } }, { "post_id": "1858713499716694264", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "So far it's a 0.25 sec hot load time on this NV dGPU laptop. Includes mapping a 4 MiB 'cart' file from pagecache, doing a 512 MiB buffer for GPU usage, hits on all PSOs, allocating some images, and kicking the command buffer that copies in the cart, and clears everything.", "timestamp": "2024-11-19 03:26:46", "media_urls": [], "reply_to_id": "1858712158004994552", "quote_of_id": null, "metrics": { "reply_count": 1, "repost_count": 0, "like_count": 0, "view_count": 98 } }, { "post_id": "1858714108415119495", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "Everything at load-time is multi-threaded to try to minimize start to in-game time. This is in sharp contrast to runtime where the only multi-threading being used is to separate things that are blocking.", "timestamp": "2024-11-19 03:29:11", "media_urls": [], "reply_to_id": "1858713499716694264", "quote_of_id": null, "metrics": { "reply_count": 1, "repost_count": 0, "like_count": 3, "view_count": 105 } }, { "post_id": "1858714908516381114", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "I get some utility on parallelizing {vulkan instance creation, mapping the cart file, warming the TLBs by walking all the pages, bringing up the window} it's about 0.08 seconds in at that point", "timestamp": "2024-11-19 03:32:22", "media_urls": [], "reply_to_id": "1858714108415119495", "quote_of_id": null, "metrics": { "reply_count": 1, "repost_count": 0, "like_count": 1, "view_count": 443 } }, { "post_id": "1858715434436227544", "author": "NOTimothyLottes", "handle": "NOTimothyLottes", "text": "After VK device is open, I signal a background thread to load the SPIR-V module, while building the descriptor set layout, which then unblocks PSO compile on background threads. And the rest of the VK setup runs in parallel. Working towards swap creation.", "timestamp": "2024-11-19 03:34:27", "media_urls": [], "reply_to_id": "1858714908516381114", "quote_of_id": null, "metrics": { "reply_count": 0, "repost_count": 0, "like_count": 2, "view_count": 424 } } ], "source_url": "https://x.com/NOTimothyLottes/status/1858715434436227544" }