Tag Archives: encoding review

Review of Multi-Resolution Encoding for HTTP Adaptive Streaming using VVenC

In their paper entitled, Multi-resolution Encoding for HTTP Adaptive Streaming using VVenC, Kamran Qureshi, Hadi Amirpour, and Christian Timmerer from the Christian Doppler Laboratory ATHENA (Alpen-Adria-Universität, Klagenfurt, Austria) propose an accelerated multi-resolution encoding strategy for HTTP Adaptive Streaming (HAS) using VVenC, an optimized open-source Versatile Video Coding (VVC) encoder. Their approach, called MEVHAS (Multi-resolution Encoding in VVenC for HTTP Adaptive …

Read More »

Sandwiched Compression: Repurposing Standard Codecs with Neural Network Wrappers

The white paper titled “Sandwiched Compression: Repurposing Standard Codecs with Neural Network Wrappers” is authored by a team from Google Research, including notable contributors such as Onur G. Guleryuz, Philip A. Chou, Berivan Isik, Hugues Hoppe, Danhang Tang, Ruofei Du, Jonathan Taylor, Philip Davidson, and Sean Fanello. Their collective expertise in image and video processing, neural networks, and rate-distortion optimization …

Read More »

Super Resolution and Content Adaptive Encoding: A Conversation with Sharon Carmel, CEO of Beamr

Jan Ozer recently caught up with Beamr’s Sharon Carmel at Mile High Video 2025. The two talked about Beamr’s latest advancements in real-time content-adaptive bitrate (CABR) optimization, super-resolution technology, and codec modernization. They also discussed Beamr’s collaboration with NVIDIA and how AI is shaping the future of video streaming. You can watch the interview on YouTube here, and it’s embedded …

Read More »

Few-Shot Domain Adaptation for Learned Image Compression

The white paper, authored by researchers from the University of Science and Technology of China, introduces a novel approach to addressing the limitations of pre-trained learned image compression (LIC) models. These models, while effective within the domains they were trained on, often falter when applied to new, domain-specific data. To overcome this, the researchers propose lightweight adapters that allow efficient …

Read More »

Cracking the Code(c): It’s All About the Implementation

Free Webinar: Cracking the Code(c): It’s All About the Implementation January 30, 2025 – 16:00 GMT/11:00 EST/08:00 PST Last fall, I published a post called “There Are No Codec Comparisons. There Are Only Codec Implementation Comparisons.” The basic premise was stated in the conclusion, which reads, “the key point is that unless the study’s testing schema and codec selection aligns …

Read More »

VMAF, SSIM, and Color Spaces: Getting the Metrics Right

So there I was, having dinner on a Saturday night date night with my wife… And in comes a text: “Hi Jan, I was wondering when you compute SSIM, do you use a single channel or multi-channel? And if multi-channel, then anything different than averaging the three?” (And yes, anyone with children, even grown children, checks incoming texts on a …

Read More »

DCVC-B: A New Deep Learning Codec for Efficient B-Frame Compression

In a recent white paper titled Bi-Directional Deep Contextual Video Compression (DCVC-B), researchers Xihua Sheng, Li Li, Dong Liu, and Shiqi Wang proposed a new approach to enhancing B-frame compression for video encoding. The authors assert that their scheme “achieves an average reduction of 26.6% in BD-Rate compared to the reference software for H.265/HEVC under random access conditions” and even …

Read More »

M3-CVC: A Glimpse into the Future of AI-Driven Video Compression

A new AI-based codec proved 18% more efficient than VVC but substantial decoding requirements will limit short-term commercial application. Here’s a summary of the white paper.  In December 2024, researchers from Fudan University introduced M3-CVC, an AI-based video compression framework that combines large multimodal models (LMMs) for semantic understanding and conditional diffusion models (CDMs) for high-fidelity reconstruction. The framework employs …

Read More »

Comparing Fixed GOPs to Variable GOPs with I-Frames at Scene Changes

I first encountered the line, “Anything worth doing is worth overdoing,” in the Robert Heinlein novel Time Enough for Love. I bring this up because this is my third recent article on GOP size, and I think I’m close to beating this topic into the ground. I’ll let you be the judge. To recount, I reported on testing in an …

Read More »

Real-World Perspectives on Choosing the Optimal GOP Size

One of the most fundamental encoding decisions is the size of the Group of Pictures (GOP) or the frequency of I-frames within an encoded file. I-frames, also known as keyframes, are the starting points for groups of pictures, consisting of I-, B-, and P-frames. Traditionally, the GOP size is directed by adaptive bitrate streaming considerations, such as ensuring an I-frame …

Read More »