Flowchart for our system.



Comparison results with other methods.


Occupancy Fields for Fast Visualizing Monocular 3D Human Representation

[基於Occupancy Field之快速視覺化三維人體表示法]
ABSTRACT:
Generating a high-quality 3D human models from a single RGB image is a challenging task. Normal maps, which record the distribution of surface normals, provide geometric features such as subtle undulations, wrinkles, and curvature variations to compensate for the limited details capture capabilities of occupancy field-based methods. Existing methods often focus on enhancing the completeness of 3D geometric structures at the expense of fast computational efficiency, thereby limiting their feasibility in interactive applications. This research introduces normal maps as auxiliary features to enhance occupancy field performance in human details reconstruction, it further combines real-time algorithms to ensure inference efficiency. We designed an optimized network architecture that effectively integrates geometric details information provided by normal maps that can improve surface texture prediction quality, ensuring a balance between detail and visual realism in reconstruction results. Experimental results demonstrated that our method outperforms existing approaches in fine geometric structure representation, texture consistency, and processing speed.

SUMMARY (中文總結):
單張 RGB 影像生成高品質的三維人體模型是一項具挑戰性的課題,廣泛應用於虛擬實境、遊戲開發與人體動作分析等領域。其中,normal map 能夠記錄物體表面的法向量分佈,提供細微起伏、皺褶與曲率變化等幾何特徵,以彌補基於 occupancy field 方法對細節捕捉能力的不足。然而,現有方法多專注於提升三維幾何結構的完整性,往往忽略了快速處理的需求,導致推論過程計算量過大,難以應用於互動式場景。本研究首先引入 normal map 作為輔助特徵,以增強 occupancy field 在細節重建上的能力,並進一步結合即時演算法,以確保推論過程的高效性。我們設計了一種優化的網路架構,使其能夠有效整合 normal map 提供的幾何細節資訊,同時提升表面紋理的預測品質,以確保重建結果在細節與視覺真實感之間取得平衡。實驗結果顯示,本方法在細微幾何結構、紋理一致性與運算速度均優於現有方法。


RESULTS:


Comparison with other methods across two datasets (THuman2.0, Multi-Garment) based on three different metrics (LPIPS, FID, hyperIQA).



REFERENCES:

    [1] Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. doi:10.1109/CVPR.2019.00459.

    [2] Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In IEEE International Conference on Computer Vision (ICCV), October 2019. doi:10.1109/ICCV.2019.00239.

    [3] Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In Proceedings of IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020. doi:10.1109/CVPR42600.2020.00016.

    [4] Ruilong Li, Yuliang Xiu, Shunsuke Saito, Zeng Huang, Kyle Olszewski, and Hao Li. Monocular real-time volumetric performance capture. In Proceedings of the European Conference on Computer Vision (ECCV), pages 49–67, 2020. doi:10.1007/978-3-030-58592-1_4.