Docs / Video-call integrations

Video-call integrations

Process the outgoing camera track so remote participants receive the beautified video. Examples for Agora, ZEGOCLOUD, LiveKit and WebRTC.

Where Facevity sits

Facevity goes between the camera capturer and the encoder of the local track. Everything downstream, the local preview and what remote participants receive, is the processed frame. Processing only the local preview is not enough, and none of the hooks below do that.

camera ─► capturer ─► [ Facevity ] ─► video source / encoder ─► network ─► remote participants
                           │
                           └─► local preview (mirrored by the renderer for the front camera)

The examples below are reference code written against each vendor’s public API for the version noted. They are not compiled against the vendor SDKs in our build, so check method names against the SDK version you ship. The underlying texture and buffer hooks are covered by the SDK’s instrumented tests.

Rules for every SDK

Topic Rule
Timestamps Pass the capture timestamp through unchanged (timestampNs). The engine never re-times frames.
Rotation Pass the frame’s rotation (degrees clockwise to upright). The output keeps the input orientation, and the RTC SDK applies its rotation metadata as before.
Mirroring Never mirror the published frame; mirror only the local preview renderer.
Formats Texture: OES or 2D in, 2D out (identity matrix). Buffers: NV21, NV12 or I420, processed in place. Use YuvPlanes for strided planes.
Backpressure processBuffer never queues. If the previous frame is still on the GPU or the budget (75% of the frame interval) is missed, the frame is filled with the last retouched frame, up to 250 ms old, and framesRepeated increments; only without a recent result does it go out unprocessed.
Threads Texture frames: call on the SDK’s GL thread and call onGlContextDestroyed() there when it stops. Buffer frames: any thread.
Cost when off fv.isActive is false when beauty is off or every level is zero. Skip conversions entirely.

Agora RTC 4.x

Register an IVideoFrameObserver at the post-capturer position with I420 frames. Agora encodes what the observer leaves in the frame, so remote users get the beautified video.

class AgoraFacevity(private val fv: Facevity) : IVideoFrameObserver {
    private var packed: ByteArray? = null

    override fun onCaptureVideoFrame(sourceType: Int, videoFrame: VideoFrame): Boolean {
        if (!fv.isActive) return true                       // nothing to do: skip the conversions
        val i420 = videoFrame.buffer.toI420()
        try {
            val w = i420.width; val h = i420.height
            val p = YuvPlanes.packI420(i420.dataY, i420.strideY, i420.dataU, i420.strideU,
                                       i420.dataV, i420.strideV, w, h, packed)
            packed = p
            val frame = BufferFrame(p, PixelFormat.I420, w, h, videoFrame.rotation,
                                    mirrored = false, timestampNs = videoFrame.timestampNs)
            if (fv.processBuffer(frame)) {                  // false = passed through (busy / budget / off)
                YuvPlanes.unpackI420(p, i420.dataY, i420.strideY, i420.dataU, i420.strideU,
                                     i420.dataV, i420.strideV, w, h)
                videoFrame.replaceBuffer(i420, videoFrame.rotation, videoFrame.timestampNs)
            }
        } finally {
            i420.release()
        }
        return true
    }

    override fun getVideoFrameProcessMode() = IVideoFrameObserver.PROCESS_MODE_READ_WRITE
    override fun getVideoFormatPreference() = IVideoFrameObserver.VIDEO_PIXEL_I420
    override fun getRotationApplied() = false    // rotation travels with the frame
    override fun getMirrorApplied() = false      // never mirror what is published
    override fun getObservedFramePosition() = IVideoFrameObserver.POSITION_POST_CAPTURER
    // onPreEncodeVideoFrame, onRenderVideoFrame, … return true (pass through)
}

// Before joining the channel:
engine.registerVideoFrameObserver(AgoraFacevity(fv))

ZEGOCLOUD Express 3.x

Enable custom video processing with GL_TEXTURE_2D before starting the preview or publishing. Zego calls the handler on its GL thread with an upright, unmirrored texture, so rotation is 0.

class ZegoFacevity(private val engine: ZegoExpressEngine, private val fv: Facevity) :
    IZegoCustomVideoProcessHandler() {

    override fun onCapturedUnprocessedTextureData(
        textureID: Int, width: Int, height: Int, referenceTimeMillisecond: Long, channel: ZegoPublishChannel,
    ) {
        val out = fv.processTexture(
            TextureFrame(textureID, false, width, height, 0, false, referenceTimeMillisecond * 1_000_000)
        )
        engine.sendCustomVideoProcessedTextureData(out, width, height, referenceTimeMillisecond, channel)
    }

    override fun onStop(channel: ZegoPublishChannel) {
        fv.onGlContextDestroyed()   // on Zego's GL thread
    }
}

// Before startPreview / startPublishingStream:
engine.enableCustomVideoProcessing(true, ZegoCustomVideoProcessConfig().apply {
    bufferType = ZegoVideoBufferType.GL_TEXTURE_2D
})
engine.setCustomVideoProcessHandler(ZegoFacevity(engine, fv))

WebRTC

Set a VideoProcessor on the video source. Camera frames arrive as OES texture buffers on the SurfaceTextureHelper thread, whose EGL context is current, so Facevity renders right there; other frames go through I420.

class WebRtcFacevity(private val fv: Facevity, private val helper: SurfaceTextureHelper) : VideoProcessor {
    private var sink: VideoSink? = null
    private var packed: ByteArray? = null
    private val yuv = YuvConverter()

    override fun setSink(sink: VideoSink?) { this.sink = sink }
    override fun onCapturerStarted(success: Boolean) = Unit
    override fun onCapturerStopped() { helper.handler.post { fv.onGlContextDestroyed() } }

    override fun onFrameCaptured(frame: VideoFrame) {
        val out = sink ?: return
        if (!fv.isActive) { out.onFrame(frame); return }
        val buffer = frame.buffer
        if (buffer is TextureBufferImpl && buffer.type == VideoFrame.TextureBuffer.Type.OES) {
            val m = RendererCommon.convertMatrixFromAndroidGraphicsMatrix(buffer.transformMatrix)
            val tex = fv.processTexture(TextureFrame(buffer.textureId, true, buffer.width, buffer.height,
                                                     frame.rotation, false, frame.timestampNs, m))
            if (tex == buffer.textureId) { out.onFrame(frame); return }
            val processed = TextureBufferImpl(buffer.width, buffer.height, VideoFrame.TextureBuffer.Type.RGB,
                tex, android.graphics.Matrix(), helper.handler, yuv, null)
            out.onFrame(VideoFrame(processed, frame.rotation, frame.timestampNs))
            return
        }
        // CPU frames (e.g. a capturer that delivers I420)
        val i420 = buffer.toI420() ?: run { out.onFrame(frame); return }
        try {
            val w = i420.width; val h = i420.height
            val p = YuvPlanes.packI420(i420.dataY, i420.strideY, i420.dataU, i420.strideU,
                                       i420.dataV, i420.strideV, w, h, packed)
            packed = p
            if (fv.processBuffer(BufferFrame(p, PixelFormat.I420, w, h, frame.rotation, false, frame.timestampNs))) {
                val next = JavaI420Buffer.allocate(w, h)
                YuvPlanes.unpackI420(p, next.dataY, next.strideY, next.dataU, next.strideU,
                                     next.dataV, next.strideV, w, h)
                out.onFrame(VideoFrame(next, frame.rotation, frame.timestampNs))
                next.release()
            } else {
                out.onFrame(frame)
            }
        } finally {
            i420.release()
        }
    }
}

videoSource.setVideoProcessor(WebRtcFacevity(fv, surfaceTextureHelper))

The texture returned by processTexture is owned by the engine and stays valid until the call after next.

LiveKit Android 2.x

LiveKit’s camera track accepts a WebRTC VideoProcessor (repackaged under livekit.org.webrtc). Reuse the WebRTC processor above with those imports and pass it when the track is created. It can’t be added later.

suspend fun publishBeautifiedCamera(room: Room, fv: Facevity, helper: SurfaceTextureHelper) {
    val local = room.localParticipant
    val track = local.createVideoTrack(
        name = "camera",
        options = LocalVideoTrackOptions(),
        videoProcessor = WebRtcFacevity(fv, helper),   // same class, livekit.org.webrtc imports
    )
    track.startCapture()
    local.publishVideoTrack(track)
}

Other SDKs

Any SDK that lets you modify captured frames before encoding works the same way: find the post-capture hook, pass the frame’s texture or YUV buffer to Facevity with its rotation and timestamp, and hand the result back. If you’re unsure where the hook is in your stack, ask us.

Need help with an integration? Contact the Facevity team. We answer integration questions during trials.