Publications

Recent Publications

Selected journal papers from the past three years are listed below. For a complete publication record, please visit Prof. Yun Zhang's personal website.

2026

Representative figure for Rate-Reconfigurable Deep Point Cloud Compression With Perceptual Bit Allocation Optimization
Rate-Reconfigurable Deep Point Cloud Compression With Perceptual Bit Allocation Optimization

Yun Zhang, Lewen Fan, Zixi Guo, Xu Wang, Xiaoxia Huang, Sam Kwong. IEEE Transactions on Image Processing, vol. 35, pp. 3451-3465, 2026.

Summary: RR-DPCC addresses the need to train separate models for different target rates by introducing a rate-reconfigurable codec with online and offline perceptual bit allocation. A single model supports flexible rate control while balancing geometry and attribute coding costs.

papercode
Representative figure for Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point Clouds
Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point Clouds

Yun Zhang, Zhiwei Guo, Zixi Guo, Linwei Zhu, C.-C. Jay Kuo. IEEE Transactions on Circuits and Systems for Video Technology, 2026.

Summary: Deep-JGAC provides end-to-end joint geometry and attribute compression for dense colored point clouds. Residual self-attention geometry coding, re-colorization, and joint entropy modeling improve the coordinated representation and coding efficiency of both components.

papercode
Representative figure for DT-JRD: Deep Transformer based Just Recognizable Difference Prediction Model for Video Coding for Machines
DT-JRD: Deep Transformer based Just Recognizable Difference Prediction Model for Video Coding for Machines

Junqi Liu, Yun Zhang, Xiaoqi Wang, Long Xu, Sam Kwong. IEEE Transactions on Multimedia, vol. 28, pp. 114-127, 2026.

Summary: The work defines JRD as the smallest difference recognizable by machine vision and formulates its prediction as multi-class classification. A Transformer jointly models content and distortion so the codec can remove unnecessary bits while preserving downstream analysis accuracy.

Representative figure for Deep Learning based Joint Geometry and Attribute Upsampling for Large-Scale Colored Point Clouds
Deep Learning based Joint Geometry and Attribute Upsampling for Large-Scale Colored Point Clouds

Yun Zhang, Feifan Chen, Na Li, Xu Wang, Fen Miao, Sam Kwong. IEEE Transactions on Image Processing, vol. 35, pp. 1305-1320, 2026.

Summary: JGAU jointly upsamples point cloud geometry and attributes by exploiting their spatial correlation. The study also establishes a large-scale colored point cloud upsampling dataset to support training and evaluation at multiple enlargement factors.

Representative figure for C-CTX: Cubic-Checkerboard Context Entropy Model for Learned Image Compression
C-CTX: Cubic-Checkerboard Context Entropy Model for Learned Image Compression

Shiyu Feng, Linwei Zhu, Yun Zhang, Na Li, Wenhui Wu, Shiqi Wang. IEEE Transactions on Multimedia, vol. 28, pp. 1756-1766, 2026.

Summary: C-CTX organizes latent context jointly across spatial and channel dimensions for learned image compression. Channel re-arrangement and multiple context prediction modes provide more accurate probability estimates and improve rate-distortion performance.

papercode
Representative figure for VP-JND: Visual Perception Assisted Deep Picture-Wise Just Noticeable Distortion Prediction Model for Image Compression
VP-JND: Visual Perception Assisted Deep Picture-Wise Just Noticeable Distortion Prediction Model for Image Compression

Yun Zhang, Shisheng Zhang, Na Li, Chunling Fan, Raouf Hamzaoui. IEEE Transactions on Circuits and Systems for Video Technology, vol. 36, no. 1, pp. 622-636, 2026.

Summary: VP-JND predicts the picture-wise distortion threshold that is just noticeable to human observers. The predicted threshold can guide coding parameters and bit allocation, reducing perceptual redundancy while maintaining visual quality.

Representative figure for RegR-PCQA: Deep Learning based Colored Point Cloud Quality Assessment Using 3D-to-2D Regularized Representation
RegR-PCQA: Deep Learning based Colored Point Cloud Quality Assessment Using 3D-to-2D Regularized Representation

Yun Zhang, Cui Mao, Na Li, Chunling Fan, Weisi Lin. IEEE Transactions on Multimedia, vol. 28, pp. 1894-1908, 2026.

Summary: The method regularizes colored point clouds into 2D geometry and attribute maps before extracting complementary cues. ViT-based geometry modeling, CNN-based attribute modeling, and quality regression are combined for objective colored point cloud assessment.

papercode
Representative figure for Temporal Consistency-Aware Dynamic Point Clouds Color Attribute Enhancement
Temporal Consistency-Aware Dynamic Point Clouds Color Attribute Enhancement

Linwei Zhu, Ruxu Liang, Yun Zhang, Hui Yuan, Sam Kwong. IEEE Transactions on Multimedia, 2026.

Summary: A 3D spatiotemporal search and temporal feature fusion pipeline addresses color noise and flicker in dynamic point clouds. Local feature extraction and Conv-PointLSTM exploit inter-frame consistency to recover more stable color attributes.

papercode

2025

Representative figure for Hierarchical Artifact Removal for Encoded Point Clouds with Very Low Bitrate G-PCC Octree
Hierarchical Artifact Removal for Encoded Point Clouds with Very Low Bitrate G-PCC Octree

Renwei Tu, Gangyi Jiang, Zhongjie Zhu, Yeyao Chen, Ting Luo, Yun Zhang, Mei Yu. IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 10, no. 1, 2025.

Summary: A hierarchical restoration method targets structural artifacts caused by very-low-bitrate G-PCC octree coding. Distorted regions are located and corrected at multiple scales to improve geometry and visual quality under sparse coding conditions.

papercode
Representative figure for Enhancing 3D Video Watching Experiences: Tackling Compression and 3D Warping Distortions in Synthesized View with Perceptual Guidance
Enhancing 3D Video Watching Experiences: Tackling Compression and 3D Warping Distortions in Synthesized View with Perceptual Guidance

Huan Zhang, Xu Zhang, Linwei Zhu, Yun Zhang, Jiangzhong Cao, Wing-Kuen Ling. Expert Systems With Applications, vol. 264, art. no. 125853, 2025.

Summary: A perceptually guided restoration network jointly addresses compression artifacts and geometric warping in synthesized 3D views. Distortion-map prediction and restoration branches cooperate to repair regions that most strongly affect the viewing experience.

papercode
Representative figure for Joint Multi-Dimensional Dynamic Attention and Transformer for General Image Restoration
Joint Multi-Dimensional Dynamic Attention and Transformer for General Image Restoration

Huan Zhang, Xu Zhang, Nian Cai, Jianglei Di, Yun Zhang. Computer Vision and Image Understanding, vol. 261, art. no. 104491, 2025.

Summary: MDDA-former combines CNNs and Transformers with multi-dimensional dynamic attention to model local texture and long-range dependencies. Its unified encoder-decoder supports general restoration tasks such as denoising, deraining, and deblurring.

Representative figure for LFIC-DRASC: Deep Light Field Image Compression Using Disentangled Representations and Asymmetrical Strip Convolution
LFIC-DRASC: Deep Light Field Image Compression Using Disentangled Representations and Asymmetrical Strip Convolution

Shiyu Feng, Yun Zhang, Linwei Zhu, Sam Kwong. IEEE Transactions on Broadcasting, vol. 71, no. 3, pp. 889-902, 2025.

Summary: LFIC-DRASC targets the high-dimensional structure and cross-view redundancy of light field imagery. Disentangled representations, asymmetrical strip convolution, and a dedicated entropy model improve compression while preserving view consistency.

Representative figure for Multi-Granular Embedding Optimization with Spatial-Channel Adaptive Tuning for Perceptual Image Quality Assessment
Multi-Granular Embedding Optimization with Spatial-Channel Adaptive Tuning for Perceptual Image Quality Assessment

Xiaoqi Wang, Yun Zhang, Junqi Liu. Neurocomputing, vol. 650, art. no. 130731, 2025.

Summary: Multi-granular auxiliary embeddings are constructed along spatial and channel dimensions to strengthen perceptual quality features. A search strategy selects effective fusion positions and orders, improving generalization across datasets and distortion types.

papercode
Representative figure for Multi-scale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines
Multi-scale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines

Junle Liu, Yun Zhang, Zixi Guo, Xiaoxia Huang, Gangyi Jiang. ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 21, no. 9, art. no. 263, 2025.

Summary: MFIBA allocates bits for multi-scale machine features according to their estimated importance. An online task loss-rate model balances feature coding cost against downstream task performance at each scale.

papercode
Representative figure for Texture-Aware Fast Mode Decision and Complexity Allocation for VVC Based Point Cloud Compression
Texture-Aware Fast Mode Decision and Complexity Allocation for VVC Based Point Cloud Compression

Lewen Fan, Yun Zhang. Journal of Visual Communication and Image Representation, vol. 113, art. no. 104610, 2025.

Summary: A texture-aware fast mode decision and complexity allocation framework reduces the cost of VVC-based point cloud video coding. Texture and coding-unit correlations prune unlikely modes early and dynamically distribute computation according to content.

papercode
Representative figure for TSC-PCAC: Voxel Transformer and Sparse Convolution-Based Point Cloud Attribute Compression for 3D Broadcasting
TSC-PCAC: Voxel Transformer and Sparse Convolution-Based Point Cloud Attribute Compression for 3D Broadcasting

Zixi Guo, Yun Zhang, Linwei Zhu, Hanli Wang, Gangyi Jiang. IEEE Transactions on Broadcasting, vol. 71, no. 1, pp. 154-166, 2025.

Summary: TSC-PCAC combines voxel Transformers and sparse convolutions for point cloud attribute compression. Its two-stage modules capture local neighborhoods and long-range dependencies, while a channel context model improves entropy coding of color attributes.

Representative figure for Geometry-Guided Latent Diffusion Model for Static Point Cloud Color Attribute Denoising
Geometry-Guided Latent Diffusion Model for Static Point Cloud Color Attribute Denoising

Linwei Zhu, Ruxu Liang, Yun Zhang, Gangyi Jiang, Yo-Sung Ho. IEEE Signal Processing Letters, vol. 32, pp. 2947-2951, 2025.

Summary: Geometry conditions a latent diffusion model for point cloud color-attribute denoising. Realistic compression-noise simulation and dynamic sampling-step control preserve structure while balancing restoration accuracy and computational efficiency.

papercode
Representative figure for Texture and Structural Distortion Metric Based on Dual-Tree Complex Wavelet Transform for DIBR-Synthesized Image Quality Assessment
Texture and Structural Distortion Metric Based on Dual-Tree Complex Wavelet Transform for DIBR-Synthesized Image Quality Assessment

Huan Zhang, Zhijun Xiong, Xu Zhang, Jiangzhong Cao, Yun Zhang. Digital Signal Processing, vol. 165, art. no. 105293, 2025.

Summary: ResNet-50 and the dual-tree complex wavelet transform extract multi-scale, multi-directional texture and structure cues. Global object-shift compensation, directional distortion statistics, and weighted regression provide targeted assessment of DIBR-synthesized images.

papercode
Representative figure for RGB-D Data Compression via Bi-Directional Cross-Modal Prior Transfer and Enhanced Entropy Modeling
RGB-D Data Compression via Bi-Directional Cross-Modal Prior Transfer and Enhanced Entropy Modeling

Yuyu Xu, Pingping Zhang, Minghui Chen, Zhuo Chen, Wenhui Wu, Yun Zhang, Xu Wang. ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 21, no. 2, art. no. 58, 2025.

Summary: The RGB-D codec establishes bi-directional cross-modal connections in the encoder, decoder, and entropy model. Bi-CPT and Bi-CEE exchange color and depth priors and use joint training to improve rate-distortion performance for both modalities.

papercode
Representative figure for Influence of Inner-Core Symmetry on Tropical Cyclone Rapid Intensification and Its Forecasting by a Machine Learning Ensemble Model
Influence of Inner-Core Symmetry on Tropical Cyclone Rapid Intensification and Its Forecasting by a Machine Learning Ensemble Model

Jiali Zhang, Qinglan Li, Liguang Wu, Qifeng Qian, Xuyang Ge, Sam Kwong, Yun Zhang, Xinyan Lyu, Guanbo Zhou, Gaozhen Nie, Pak Wai Chan, Wai Kin Wong, Linwei Zhu. Weather and Climate Extremes, vol. 48, art. no. 100770, 2025.

Summary: The study links tropical cyclone inner-core symmetry with rapid intensification and builds multiple tree-based predictors plus an ensemble model. Satellite imagery, intensity, and environmental variables are fused to support objective forecasts at different lead times.

papercode

2024

Representative figure for Colored Point Cloud Quality Assessment Using Complementary Features in 3D and 2D Spaces
Colored Point Cloud Quality Assessment Using Complementary Features in 3D and 2D Spaces

Mao Cui, Yun Zhang, Chunling Fan, Raouf Hamzaoui, Qinglan Li. IEEE Transactions on Multimedia, vol. 26, pp. 11111-11125, 2024.

Summary: CF-PCQA treats local structure in 3D point space and texture in 2D projections as complementary evidence. Multi-branch feature extraction and regression jointly learn how geometry, color, and spatial distribution affect perceived point cloud quality.

papercode
Representative figure for Learning to Predict Object-Wise Just Recognizable Distortion for Image and Video Compression
Learning to Predict Object-Wise Just Recognizable Distortion for Image and Video Compression

Yun Zhang, Haoqing Lin, Jing Sun, Linwei Zhu, Sam Kwong. IEEE Transactions on Multimedia, vol. 26, pp. 5925-5938, 2024.

Summary: The work refines just recognizable distortion from the image level to individual objects and constructs a corresponding data and prediction pipeline. Region-specific thresholds provide finer perceptual guidance for content-adaptive image and video compression.

Representative figure for Neural Network Based Multi-Level In-Loop Filtering for Versatile Video Coding
Neural Network Based Multi-Level In-Loop Filtering for Versatile Video Coding

Linwei Zhu, Yun Zhang, Na Li, Wenhui Wu, Shiqi Wang, Sam Kwong. IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 11, pp. 12092-12096, 2024.

Summary: Three neural filters are inserted at reference-pixel, coding-unit, and frame levels inside the VVC loop. They address local prediction error, block distortion, and global artifacts respectively, and their combination further improves reconstruction and coding efficiency.

papercode
Representative figure for Quality Assessment for DIBR-Synthesized Views Based on Wavelet Transform and Gradient Magnitude Similarity
Quality Assessment for DIBR-Synthesized Views Based on Wavelet Transform and Gradient Magnitude Similarity

Huan Zhang, Dongsheng Zheng, Yun Zhang, Jianzhong Cao, Weisi Lin, Wing-Kuen Ling. IEEE Transactions on Multimedia, vol. 26, pp. 6834-6847, 2024.

Summary: Wavelet decomposition and gradient-magnitude similarity capture black holes, stretching, and structural displacement in DIBR-synthesized views. Reference-view compensation and multi-scale fusion increase sensitivity to localized geometric distortion.

papercode
Representative figure for Dynamic Weighted Gradient Reversal Network for Visible-Infrared Person Re-Identification
Dynamic Weighted Gradient Reversal Network for Visible-Infrared Person Re-Identification

Chenghua Li, Zongze Li, Jing Sun, Yun Zhang, Xiaoping Jiang, Fan Zhang. ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 20, no. 1, art. no. 12, 2024.

Summary: DGRNet applies dynamically weighted gradient reversal to visible-infrared person re-identification. Identity classification and modality discrimination are trained jointly, with adaptive adversarial strength reducing the modality gap while preserving identity cues.

papercode
Representative figure for Multiview Projection Based Joint Geometry and Color Hole Repairing Method for G-PCC Trisoup Encoded Color Point Cloud
Multiview Projection Based Joint Geometry and Color Hole Repairing Method for G-PCC Trisoup Encoded Color Point Cloud

Taowen Xu, Gangyi Jiang, Mei Yu, Yun Zhang, Zhidi Jiang, Yo-Sung Ho. IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 8, no. 1, pp. 892-902, 2024.

Summary: A multi-view projection framework jointly repairs geometry holes and missing colors introduced by G-PCC Trisoup coding. Geometry-weighted repair and a gated color network work together before inverse projection reconstructs a more complete colored point cloud.

papercode
Representative figure for PCQD-AR: Subjective Quality Assessment of Compressed Point Clouds with Head-Mounted Augmented Reality
PCQD-AR: Subjective Quality Assessment of Compressed Point Clouds with Head-Mounted Augmented Reality

Chunling Fan, Yun Zhang, Linwei Zhu, Xinju Wu. IET Electronics Letters, vol. 60, no. 5, pp. 1-4, 2024.

Summary: A head-mounted AR subjective study builds a dataset containing 10 reference point clouds and 90 distorted versions. It analyzes how geometry and texture quantization affect six-degree-of-freedom viewing and provides a benchmark for objective quality metrics.

papercode
Representative figure for Cracks-Suppression Perceptual Geometry Coding for Dynamic Point Clouds
Cracks-Suppression Perceptual Geometry Coding for Dynamic Point Clouds

Wei Liu, Mei Yu, Zhidi Jiang, Haiyong Xu, Zhoujie Zhu, Yun Zhang, Gangyi Jiang. Digital Signal Processing, vol. 149, pp. 1-13, 2024.

Summary: Perceptual geometry coding targets cracks introduced during V-PCC projection, filling, and quantization. Improved geometry-map generation and coding control preserve continuity in visually important regions of reconstructed dynamic point clouds.

papercode
Representative figure for Optimized Quantization Parameter Selection for Video-Based Point Cloud Compression
Optimized Quantization Parameter Selection for Video-Based Point Cloud Compression

Hui Yuan, Raouf Hamzaoui, Ferrante Neri, Shengxiang Yang, Xin Lu, Linwei Zhu, Yun Zhang. Frontiers in Signal Processing, vol. 4, art. no. 1385287, 2024.

Summary: Rate-distortion optimization and differential evolution select V-PCC quantization parameters for groups of frames. Content- and rate-aware parameter search reduces bitrate while maintaining reconstructed point cloud quality.

papercode
Representative figure for Guest Editorial: Deep Learning-Based Point Cloud Processing, Compression and Analysis
Guest Editorial: Deep Learning-Based Point Cloud Processing, Compression and Analysis

Yun Zhang, Raouf Hamzaoui, Xu Wang, Junhui Hou, Giuseppe Valenzise. IET Electronics Letters, vol. 60, no. 14, pp. 1-2, 2024.

Summary: This guest editorial reviews recent deep-learning advances in point cloud processing, compression, and analysis. It covers sampling, enhancement, coding, quality assessment, and 3D understanding while summarizing the main contributions of the special issue.

papercode

Papers

More than 180 SCI/EI-indexed papers, including over 70 papers in IEEE/ACM Transactions.

Monograph

Author of the academic monograph Three-Dimensional Video Signal Processing.

Patents

More than 50 granted Chinese, U.S., and PCT invention patents.