Motion detection technology superimposes expressions onto target faces in real time.
The first-order motion model takes a slightly different approach by replacing the encoder with a motion model. The underlying AI detects facial expressions, eye movements, and head position, which are then superimposed on a destination image. Anyone who has used Snap will have used this approach. The neural network is trained on many hours of real video footage to recognize various important features of a person’s face. Since a video is a collection of images or frames, this allows photoshopping each frame in a video.
— Overview Of How To Create Deepfakes - It’s Scarily Simple · Forbes