Skip to content

Performance: CPU and GPU Kernel Optimization Opportunities #1676

Description

@simonCatBot

Performance: CPU and GPU Kernel Optimization Opportunities

Summary

Comprehensive benchmarking on AMD Ryzen AI MAX+ PRO 395 (32 cores) with Radeon 8060S GPU has identified specific OpenVX kernels with significant performance improvement opportunities. This issue provides a prioritized list of kernels that would benefit from optimization on both CPU and GPU backends.

System Information

  • CPU: AMD Ryzen AI MAX+ PRO 395 (32 cores/64 threads)
  • GPU: AMD Radeon 8060S (40 CUs)
  • RAM: 94GB
  • OS: Ubuntu 24.04, Linux 6.17.0-35-generic
  • MIVisionX Version: 3.5.0
  • OpenVX Version: 1.3
  • Benchmark: 100 iterations per kernel across VGA, FHD, and 4K resolutions

CPU Backend Optimization Candidates

These kernels show performance gaps indicating opportunities for AMD CPU optimization (AVX2/AVX-512, better memory access patterns, or improved algorithms).

🔴 Critical Priority (Severe Performance Issues)

These kernels achieve less than 50% of expected throughput and likely have implementation issues:

  • LaplacianPyramid_S16

    • Status: 0.2-4.5% of expected throughput
    • Likely Cause: Unimplemented or placeholder scalar implementation; missing SIMD path for S16
  • ScaleImage_Area_Half

    • Status: 31-81% of expected throughput
    • Likely Cause: Suboptimal memory access patterns for area-based accumulation
  • NonLinearFilter_Max

    • Status: 36-72% of expected throughput
    • Likely Cause: Generic implementation without architecture-specific optimizations
  • NonLinearFilter_Min

    • Status: 37-72% of expected throughput
    • Likely Cause: Generic implementation without architecture-specific optimizations
  • CustomConvolution_U8_S16

    • Status: ~45% of expected throughput (consistent across resolutions)
    • Likely Cause: S16 output path lacks SIMD optimization
  • Threshold_Range

    • Status: 46% of expected throughput at VGA
    • Likely Cause: Range check overhead not optimized

🟡 High Priority (Significant Performance Gaps)

These kernels achieve 50-85% of expected throughput:

Color Space Conversions

  • ColorConvert_IYUV2RGB (60-66% of expected throughput)
  • ColorConvert_RGB2YUV4 (66-67% of expected throughput)
  • ColorConvert_RGB2IYUV (81-83% of expected throughput)
  • Likely Cause: Missing hand-optimized SIMD implementations for color conversion matrices

S16 Arithmetic Operations

  • Multiply_S16_S16_S16 (70-74% of expected throughput)
  • Add_S16_S16_S16 (82% of expected throughput)
  • Likely Cause: Incomplete SIMD coverage for 16-bit signed integer operations

Feature Detection

  • FastCorners (84-85% of expected throughput)
  • HarrisCorners (83-94% of expected throughput)
  • HarrisTracker (83-94% of expected throughput)
  • Likely Cause: Generic corner detection without architecture-specific optimizations

🟠 Medium Priority (Moderate Improvements Possible)

These kernels achieve 70-90% of expected throughput:

  • And_U8_U8_U8 / Or_U8_U8_U8 / Xor_U8_U8_U8 / Not_U8_U8 (67-89%)
  • Histogram (~73% of expected throughput)
  • HistogramEqualize (~73% of expected throughput)
  • IntegralImage (78-92% of expected throughput)
  • MeanStdDev (~67% of expected throughput at VGA)

GPU Backend Optimization Candidates

These kernels fail to achieve the 10x speedup target over CPU performance on AMD Radeon 8060S GPU at 4K resolution.

🔴 Critical Priority (< 2x Speedup - Poor GPU Utilization)

  • OpticalFlowPyrLK

    • GPU Speedup: 1.01x
    • Status: Effectively no GPU acceleration
  • WarpPerspective

    • GPU Speedup: 1.84x
    • Status: Far below 10x target; needs investigation
  • Remap_Nearest

    • GPU Speedup: 2.02x
    • Status: Minimal GPU benefit
  • Remap

    • GPU Speedup: 2.23x
    • Status: Needs optimization for parallel interpolation
  • ChannelExtract

    • GPU Speedup: 2.45x
    • Status: Simple operation should achieve higher parallelism

🟡 High Priority (2-5x Speedup - Below Target)

  • ScaleImage_Half

    • GPU Speedup: 2.58x
    • Status: Memory-bound operation, consider shared memory optimization
  • LaplacianPyramid

    • GPU Speedup: 1.59x
    • Status: Multi-scale decomposition needs kernel fusion
  • LaplacianReconstruct_S16

    • GPU Speedup: 4.69x
    • Status: Close to target but S16 path may have issues
  • CannyEdgeDetector

    • GPU Speedup: 5.09x at 4K (fails at FHD due to memory)
    • Status: Complex algorithm needs staged optimization
  • ChannelExtract_YUYV_Y

    • GPU Speedup: 6.17x
    • Status: Channel extraction should achieve higher throughput
  • ColorConvert_IYUV2RGB

    • GPU Speedup: 6.58x
    • Status: Color matrix multiplication optimizable
  • Box3x3

    • GPU Speedup: 7.86x
    • Status: Simple filter should exceed 10x
  • ScaleImage_Double

    • GPU Speedup: 8.95x
    • Status: Nearly at target
  • Add_U8_U8_S16

    • GPU Speedup: 9.0x (at FHD, 24x at 4K)
    • Status: Approaches target at higher resolutions

🟠 Medium Priority (5-10x Speedup - Near Target)

  • Gaussian3x3 (3.76x)
  • Median3x3 (3.62x)
  • EdgeDetection (3.63x - fails at FHD)
  • HalfScaleGaussian (3.63x)
  • DualFilter (3.97x)
  • MeanStdDev (2.19x)
  • MinMaxLoc_S16 (2.14x)
  • Sobel3x3 (1.37x)
  • Multiply_U8_U8_S16 (1.69x)
  • Erode3x3 (1.46x)
  • Dilate3x3 (expected similar to Erode)

⚠️ Failed GPU Kernels (Memory Issues at High Resolution)

These kernels fail at FHD/4K due to GPU memory allocation:

  • CannyEdgeDetector (hipMemcpyDtoH failed at FHD+)
  • EdgeDetection (hipMemcpyDtoH failed at FHD+)

Root Cause: Intermediate buffer allocation exceeds GPU memory at high resolution.


GPU Performance Success Stories (Reference)

These kernels demonstrate the performance potential and serve as implementation references:

Kernel 4K Speedup Notes
WarpAffine 56.8x Excellent parallelization of geometric transforms
Subtract_U8_U8_S16 37.2x Simple pixel-wise operations achieve massive speedup
Add_U8_U8_S16 24.0x U8→S16 arithmetic highly optimized
TableLookup 55x Memory-bound but well-optimized
WarpPerspective_Nearest 102x Nearest neighbor interpolation efficient on GPU
CustomConvolution 40.7x Convolution with proper tiling
ColorConvert_RGB2YUV4 22-56x Color conversion can achieve high throughput

Resolution Scaling Insights

CPU Backend

  • Performance degrades 5-15% from VGA to 4K
  • Sub-linear scaling due to memory bandwidth limitations
  • Cache efficiency drops with larger image sizes

GPU Backend

  • Performance increases 100-500% from VGA to 4K
  • Super-linear scaling due to better parallelism utilization
  • Fixed overhead (kernel launch, memory setup) amortized over more pixels
  • Key Finding: Many kernels below 10x at VGA achieve >10x at 4K

Recommendations by Priority

Immediate Action (Critical Priority)

  1. LaplacianPyramid_S16 - Verify implementation completeness
  2. OpticalFlowPyrLK - Investigate why GPU shows no speedup
  3. ScaleImage_Area_Half - Optimize memory access patterns
  4. NonLinearFilter variants - Implement specialized SIMD paths

Short Term (High Priority)

  1. Color space conversion SIMD implementations
  2. S16 arithmetic operation vectorization
  3. WarpPerspective GPU kernel optimization
  4. Simple filters (Box3x3, Gaussian3x3, Sobel3x3) GPU tuning

Medium Term (Medium Priority)

  1. Feature detection algorithm optimization
  2. Histogram and statistical operations
  3. Channel extraction GPU kernels
  4. Morphological operations (Erode/Dilate)

Benchmark Methodology

  • Iterations: 100 per kernel (10 warmup)
  • Resolutions: VGA (640×480), FHD (1920×1080), 4K (3840×2160)
  • CPU Threading: OpenMP with 32 threads
  • Metrics: Megapixels/second throughput
  • Statistical Significance: CV% < 5%
  • GPU: AMD Radeon 8060S with ROCm/HIP backend

Attachments

Full benchmark data and profiling results available upon request. The data shows clear patterns for which kernel categories would benefit most from:

  • S16 data type SIMD implementations (CPU)
  • Memory access pattern optimization (CPU/GPU)
  • Kernel fusion for multi-scale operations (GPU)
  • Shared memory utilization for filter operations (GPU)

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions