✓ Indexed in BIBNEX
Info:eu Repo/semantics/article
Comparative Analysis of Deep fake Video and Audio Detection
Abstract
Deepfake ("deep learning + fake" = DF) refers to the forged videos and audios generated using AI algorithms. While they can be a source of entertainment,theycanalsobeharmfulinvariousways.Manipulatingbothau- diosandvideosforharmfulpurposeshasbeen a concerning issue from the past more than 10 years. The ability todetect these videos and audios through AI detectors is a motivatingfactor in achieving the best results for the project. Thispapercontainscomparativestudyoftheexistingresearchondeepfakeissues showcasingAccuracy,F1Score,Barplotsandgraphsofthesame.Whileexplo- rationofdeepfakevideoshaveseveralapproachesanddatasetsavailable,theau- diodeepfakes has been relatively neglected. In this work, we propose theidea ofjointdeepfakevideoandaudiodetectionusingahybrid deep learning model ensembling ResNet50 and EfficientNet B0. The dataset comprising real and synthetic voice recordings was selected from the SceneFake repository on Kaggle. Key audio features, including Mel- frequency cepstral coefficients (MFCCs), spectral centroid, chroma, zero- crossing rate, and root mean square energy(RMSE),wereextractedusingthelibrosalibrary. RandomForestClassi- fierwastrainedtodetectaudiowhileDFDVdatasetwasutilizedtoextractfacial frames from videos.
Keywords
Deep fake Detection
ResNet50
Efficient Net B0
MFCC
Random Forest Classifier
Citation
Comparative Analysis of Deep fake Video and Audio Detection.
Journal of Recent Innovation in Science and Technology
.
2025.
Vol. 1
(1)
DOI: 10.70454/jrist.2025.10102