% ======================================================================
%  Audio-JEPA Representations for Speech Deepfake Detection
%  SOICT 2026 — Springer CCIS — llncs
%  Biên dịch: pdflatex main && bibtex main && pdflatex main && pdflatex main
%  Ghi chú tiếng Việt cho tác giả nằm trong comment "%% VN:".
% ======================================================================
%% VN: SOICT yêu cầu bản nộp KHÔNG có số trang. Bỏ option runningheads thì llncs
%%     thật chọn \ps@empty, tức tắt cả số trang lẫn running head. Giữ nguyên
%%     \titlerunning và \authorrunning bên dưới để bật lại cho camera-ready.
\documentclass{llncs}

\usepackage[T1]{fontenc}
\usepackage[utf8]{inputenc}
% template LNCS 2026 (llncs v2.25) dùng Times qua newtx. Sandbox không có newtx nên
% dự phòng mathptmx; trên Overleaf nhánh newtx được chọn.
\IfFileExists{newtxtext.sty}{\usepackage{newtxtext}\usepackage[varvw]{newtxmath}}{\usepackage{mathptmx}}
\usepackage{graphicx}
\usepackage{booktabs}
\usepackage{array}
\setlength{\tabcolsep}{4pt}   % llncs mặc định 1.4pt, quá sát
\usepackage{multirow}
\usepackage{amsmath}
\usepackage{url}
\usepackage{xcolor}
\usepackage[hidelinks]{hyperref}
\renewcommand\UrlFont{\color{blue}\rmfamily}

% số liệu hay dùng — đổi một chỗ, đổi cả bài
\newcommand{\eer}{EER}
\newcommand{\pct}{\,\%}
\newcommand{\TODO}[1]{\textcolor{red}{[TODO: #1]}}

\begin{document}

\title{Audio-JEPA Representations for Speech Deepfake Detection}
\titlerunning{Audio-JEPA for Speech Deepfake Detection}

%% VN: Giữ cách viết tên không dấu do tác giả cung cấp; không phải giới hạn bắt buộc của LaTeX.
%%     Không chèn ORCID trong bản này theo quyết định của tác giả.
%% VN: dạng khai đơn vị lấy theo chính bài của thầy Hoài trên Springer —
%%     FDSE 2024 (CCIS 2309, cùng series SOICT) và ICIT 2024 (LNDECT 230):
%%     Khoa KH&KTMT + HCMUT, rồi ĐHQG-HCM, mỗi tác giả mang cả hai số mũ.
\author{Nguyen Le Nguyen\inst{1,2} \and Tran Van Hoai\inst{1,2}}
\authorrunning{N. L. Nguyen and V. H. Tran}
\institute{Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Ho Chi Minh City, Vietnam \and
Vietnam National University Ho Chi Minh City (VNU-HCM), Ho Chi Minh City, Vietnam\\
\email{\{nlnguyen.sdh231,hoai\}@hcmut.edu.vn}}

\maketitle

\begin{abstract}
Speech deepfake detectors are built on self-supervised front-ends, but published systems differ in architecture, size and pre-training corpus at once, so the contribution of the pre-training objective itself is hard to read off. We apply Audio-JEPA --- a ViT-Base encoder pre-trained on AudioSet to predict latent representations of masked spectrogram patches --- combining its layer outputs by learned weights, attentive statistics pooling and a small classification head. We train once on the ASVspoof2019 LA training partition with a fixed recipe; evaluation runs on the disjoint evaluation partition and further corpora through the Speech DF Arena toolkit, against verified published baselines. Under an identical downstream recipe we compare it against AudioMAE, a released checkpoint that shares its ViT-Base backbone and its AudioSet pre-training corpus but is trained to reconstruct masked patches rather than to predict them: latent prediction is ahead in three of the four comparisons, by 2.57 to 7.55 \eer{} points over three seeds per arm with no overlap between seed ranges. With the encoder frozen, the system reaches 8.47\pct{} \eer{} on DFADD, where our re-run of the AASIST baseline gives 39.05\pct{} and the best system trained on the same corpus, the 340\,M-parameter XLSR+SLS, gives 7.54\pct. Accuracy depends on how closely evaluation data matches the training corpus; we analyse that dependence and what it implies for the choice of pre-training corpus.

\keywords{speech deepfake detection \and speech anti-spoofing \and self-supervised learning \and ASVspoof \and representation learning \and joint-embedding predictive architecture \and cross-corpus generalization}
\end{abstract}

% ======================================================================
\section{Introduction}
\label{sec:intro}

Detecting synthetic or converted speech --- the \emph{logical access} spoofing task of the ASVspoof challenges~\cite{todisco2019asvspoof} --- has become a problem of generalisation rather than of in-domain accuracy. On the ASVspoof2019 LA evaluation partition, ten of the twelve open-source systems collected by the Speech DF Arena benchmark~\cite{dowerah2025arena} are below 5\,\% equal error rate and the best is at 0.12\,\%, but on corpora recorded under other conditions or attacked by newer generators the same systems range from a few percent to near chance. Every system near the top of that leaderboard is built on a self-supervised speech front-end: wav2vec 2.0 and XLS-R~\cite{baevski2020wav2vec2,babu2022xlsr}, trained with a contrastive objective on hundreds of thousands of hours of speech, or WavLM and HuBERT~\cite{chen2022wavlm,hsu2021hubert}, trained to predict masked discrete units.

One option is absent from that leaderboard: a pre-training target that is continuous and produced by the model itself. Joint-embedding predictive architectures~\cite{lecun2022path,assran2023ijepa,bardes2024vjepa} learn by predicting the \emph{representation} of a masked region from its context, with the target supplied by an exponential moving average of the encoder being trained rather than by a codebook, a set of negatives or the input signal; Audio-JEPA~\cite{tuncay2025audiojepa} recently applied the recipe to spectrograms and released a ViT-Base encoder pre-trained on AudioSet. Among the twelve open-source systems in the Speech DF Arena paper, no front-end is pre-trained this way, and none is pre-trained to reconstruct its input either; the closest prior work we know of evaluated AudioMAE features for deepfake detection and found them the weakest of the deep features it compared, attributing the result to the AudioSet pre-training corpus~\cite{yang2024features}.

Which property of a front-end matters is hard to read off the leaderboard, because the systems on it differ in architecture, parameter count, pre-training corpus and input representation at the same time. This paper applies Audio-JEPA to speech deepfake detection with a light back-end, trained once on ASVspoof2019 LA with a fixed recipe and evaluated on five public corpora through the Speech DF Arena toolkit, after checking that toolkit against the published baseline numbers. It then places that result beside AudioMAE~\cite{huang2022audiomae}, the closest available checkpoint --- same backbone, same pre-training corpus, a reconstruction objective instead of a predictive one --- run through an identical downstream recipe.

Latent prediction leads masked reconstruction by 2.57 to 7.55 EER points in three of the four comparisons we ran, across two corpora and both adaptation regimes, with no overlap between seed ranges.

% ======================================================================
\section{Background}
\label{sec:background}

\paragraph{Self-supervised front-ends for anti-spoofing.} Tak et al.~\cite{tak2022w2v2aasist} showed that a fine-tuned wav2vec 2.0 XLS-R encoder feeding the AASIST graph-attention back-end~\cite{jung2022aasist}, trained with RawBoost augmentation, generalised far better than end-to-end models trained from scratch; Wang and Yamagishi~\cite{wang2022ssl_frontends} compared several speech SSL front-ends under one back-end and reached the same conclusion. The systems that lead the Speech DF Arena leaderboard among those trained on ASVspoof2019 --- XLSR+SLS~\cite{zhang2024xlsrsls}, TCM~\cite{truong2024tcm}, XLSR-Mamba~\cite{xiao2025xlsrmamba}, Nes2Net~\cite{liu2025nes2net} --- all keep an XLS-R front-end of about 300\,M parameters and vary the back-end. WavLM, HuBERT and wav2vec 2.0 encoders with an ECAPA-TDNN~\cite{desplanques2020ecapa} back-end complete the set. All of these front-ends are speech models, pre-trained on speech corpora ranging from the 960 hours of LibriSpeech for the base HuBERT and WavLM checkpoints to the 436\,000 hours behind XLS-R.

\paragraph{Where the pre-training target comes from.} Self-supervised objectives can be told apart by what they ask the encoder to predict and by what supplies the answer. \emph{Contrastive} objectives (wav2vec 2.0) identify the true quantised target of a masked frame among distractors drawn from the same utterance. \emph{Masked-unit prediction} (HuBERT, WavLM) predicts cluster identities assigned offline by $k$-means. \emph{Masked reconstruction} (AudioMAE~\cite{huang2022audiomae}) decodes masked spectrogram patches back to their values and minimises the error in input space. \emph{Latent prediction} (I-JEPA, Audio-JEPA) uses continuous targets supplied by an exponential moving average of the encoder: with a context encoder $f_\theta$, a predictor $g_\phi$ and a target encoder $f_{\bar\theta}$ updated as an exponential moving average of $f_\theta$, it minimises
\begin{equation}
\mathcal{L} = \big\| g_\phi\big(f_\theta(x_{\mathrm{ctx}}), m\big) - \mathrm{sg}\,\big[f_{\bar\theta}(x)\big]_{M} \big\|_2^2 ,
\end{equation}
where $m$ encodes the positions of the masked target block and $\mathrm{sg}$ stops gradients. The target encoder processes the full input $x$; its outputs at the masked positions $M$ provide the targets. The representation is free to drop whatever the predictor cannot predict; whether that helps or hurts for detecting synthesis artefacts is an empirical question, and the one this paper asks.

\paragraph{Speech DF Arena.} Dowerah et al.~\cite{dowerah2025arena} released an evaluation toolkit and a leaderboard covering fourteen deepfake corpora. Their paper reports twelve open-source systems, all trained on ASVspoof2019 LA (Whisper-MesoNet~\cite{kawa2023whisper} excepted), with released per-utterance score files; this is what makes an independent re-run checkable. We evaluate on five of the fourteen corpora (Sect.~\ref{sec:datasets}), chosen to span genuine-speech sources and spoofing types.

% ======================================================================
\section{Method}
\label{sec:method}

\subsection{Front-end: Audio-JEPA}
\label{sec:frontend}

Audio-JEPA~\cite{tuncay2025audiojepa} applies the joint-embedding predictive architecture of I-JEPA~\cite{assran2023ijepa} to log-mel spectrograms. During pre-training a context encoder sees a subset of spectrogram patches, and a small predictor is trained to output the representations that a target encoder --- an exponential moving average of the context encoder --- assigns to masked blocks of patches. The loss is computed in representation space; no patch is ever reconstructed. The published checkpoint\footnote{\url{https://huggingface.co/ltuncay/Audio-JEPA}} follows the ViT-Base configuration~\cite{dosovitskiy2021vit} (12 blocks, width 768, 12 heads, 85.4\,M parameters) and is pre-trained on AudioSet-2M~\cite{gemmeke2017audioset}: 5\,338 hours of general audio spanning speech, music and environmental sound~\cite{tuncay2025audiojepa}. We use the context encoder and do not use the target encoder or the predictor.\footnote{Both I-JEPA and Audio-JEPA evaluate with the target encoder, which is an exponential moving average of the context encoder. We use Audio-JEPA's context encoder; the effect of choosing it rather than its target encoder was not measured.}

The toolkit hands the model a 16\,kHz waveform tiled or truncated to 64\,600 samples (Sect.~\ref{sec:toolkit}). We resample to 32\,kHz, following the Audio-JEPA convention --- this adds no information, since the source is 16\,kHz --- take the first 2.56\,s (tiling if shorter), and compute a 128-band log-mel spectrogram with a 25\,ms window, 10\,ms hop and a 20--8\,000\,Hz range, with per-utterance mean subtraction, zero-padded or truncated to 256 frames. The $256\times128$ spectrogram is cut into patches of 8 frames by 32 mel bins (80\,ms $\times$ a quarter of the mel axis), giving a grid of $32\times4=128$ tokens, each projected linearly to 768 dimensions and given a two-dimensional sinusoidal position embedding; there is no class token. The pre-trained patch embedding is $16\times16$; we interpolate its weights bicubically to $8\times32$ and regenerate the position embeddings for the new grid. The token count is unchanged by this: Audio-JEPA's $16\times16$ patches over a $128\times256$ spectrogram give 128 patches, and our $(8,32)$ patches over $256\times128$ give $32\times4$, also 128. These counts refer to the full patch grids before masking.

These input settings differ from the pre-training recipe, which uses 10\,s clips at 32\,kHz with 128 mel bands over 256 frames~\cite{tuncay2025audiojepa} --- about a 39\,ms hop, four times coarser in time than ours, with the released configuration setting the upper frequency to 16\,kHz.\footnote{\url{https://github.com/LudovicTuncay/Audio-JEPA}} That configuration is tuned for sound-event tagging at a ten-second scale; when we ran it for this task the held-out attack improved fivefold while the evaluation partition got worse, so we kept the finer resolution.

\subsection{Back-end}
\label{sec:backend}

The back-end is deliberately conventional. Learned layer weighting is the standard way to read a self-supervised encoder~\cite{wang2022ssl_frontends}, and attentive statistics pooling~\cite{okabe2018asp} is the pooling used inside the ECAPA-TDNN back-ends on the same leaderboard. Both arms of Sect.~\ref{sec:matched} use this back-end unchanged.

The encoder yields thirteen token sequences of shape $128\times768$: the output of each of the twelve blocks and the final layer-normalised output. Three steps turn them into a score. A learned softmax over thirteen scalars weights the sequences and sums them (13 parameters, initialised to zero and hence uniform after the softmax). Attentive statistics pooling~\cite{okabe2018asp} --- a two-layer attention network ($768\to128\to1$, tanh) with a softmax over the 128 tokens --- produces attention-weighted mean and standard deviation vectors, concatenated to 1\,536 dimensions (98\,561 parameters). A classifier of LayerNorm, Linear $1\,536\to256$, GELU, Dropout 0.2 and Linear $256\to2$ gives two logits (397\,058 parameters). The back-end totals 495\,632 parameters, about 0.5\,M, or 0.58\,\% of the encoder. The score written for each utterance is the logit difference $d = z_{\mathrm{bonafide}} - z_{\mathrm{spoof}}$; we use the difference rather than a single logit because the softmax probability is monotone in the difference, not in either logit alone, and the toolkit's \eer{} depends only on score order.


Figure~\ref{fig:pipeline} distinguishes the published pre-training procedure from the downstream pipeline used here.

\begin{figure}[!t]
\centering
\includegraphics[width=\linewidth]{figures/pipeline.pdf}
\caption{(a) Audio-JEPA pre-training, schematically redrawn from~\cite{tuncay2025audiojepa}: the context branch predicts representations at masked positions; the target encoder sees the full input and is updated by exponential moving average (EMA). Gradients do not pass through the target branch. (b) Downstream detection using the released context encoder; we do not repeat pre-training. Its twelve block outputs and final layer-normalised output are aggregated. Only the back-end is trained in the frozen regime; both encoder and back-end are trained in the fine-tuned regime.}
\label{fig:pipeline}
\end{figure}

\subsection{Training}
\label{sec:training}

We train in two regimes. \emph{Frozen}: the encoder runs in evaluation mode without gradients and only the back-end is trained. \emph{Fine-tuned}: all parameters are trained, with a learning rate of $3\times10^{-5}$ for the encoder and $10^{-3}$ for the back-end. Both use AdamW~\cite{loshchilov2019adamw} with weight decay 0.05, a one-cycle schedule with 10\,\% warm-up, batch size 32, mixed precision and exactly six epochs. The loss is cross-entropy with the bona fide class up-weighted by the spoof-to-bona-fide ratio of the training set (about 9.6). The only augmentation, applied during training, is a SpecAugment-style~\cite{park2019specaugment} mask drawn afresh for each spectrogram: one band of up to 16 of the 256 frames and one of up to 16 of the 128 mel bins, at random positions --- at most 6.25\,\% of the time axis and 12.5\,\% of the mel axis. This is unrelated to the patch masking used to pre-train the encoder, which covers 40--60\,\% of patches and defines the learning task rather than regularising it. Training data and hold-out are described in Sect.~\ref{sec:datasets}; the final six-epoch checkpoint is used for each main configuration (Sect.~\ref{sec:toolkit}).

Every configuration in Tables~\ref{tab:main} and~\ref{tab:objective} is trained with three seeds, 1234, 2 and 3, except where a single seed is marked. We report both the frozen and the fine-tuned regime, not to choose between them but because they give that comparison a second axis: a difference that survives with the encoder frozen cannot be an artefact of how fine-tuning interacts with one checkpoint or the other.

\subsection{Comparison against a reconstruction-trained checkpoint}
\label{sec:matched}

The closest available point of comparison for Audio-JEPA is AudioMAE~\cite{huang2022audiomae}: the same ViT-Base backbone, the same AudioSet-2M pre-training corpus, and a reconstruction objective in place of a predictive one. We take the encoder of the public checkpoint (the \texttt{timm} port \texttt{vit\_base\_patch16\_\allowbreak 1024\_128.\allowbreak audiomae\_as2m}), interpolate its patch embedding to $8\times32$ in the same way, and give it the same input, back-end, training data, recipe and seeds.

What this holds fixed is the backbone, the pre-training corpus and everything downstream; what it does not hold fixed is the rest of each checkpoint's pre-training recipe, which the two papers set independently --- masking ratio, sample rate, optimiser and compute budget among them. The two encoders are also displaced from their pre-training input configuration along different axes: our 2.56\,s input shortens the clip about fourfold for both, preserves Audio-JEPA's full patch count while quartering AudioMAE's, and moves Audio-JEPA's patch from 625 to 80\,ms against AudioMAE's 160 to 80\,ms. No single configuration is native to both, and giving each its own would change the input alongside the objective. We therefore read the comparison as one between two released checkpoints under a common downstream recipe, and treat the attribution of the difference to the objective itself as a hypothesis, examined in Sect.~\ref{sec:discussion} and bounded in Sect.~\ref{sec:limitations}.

% ======================================================================
\section{Experimental Setup}
\label{sec:setup}

\subsection{Evaluation toolkit and protocol}
\label{sec:toolkit}

All evaluations use the Speech DF Arena toolkit~\cite{dowerah2025arena}. In the fixed-length configuration we run, the toolkit loads every file at 16\,kHz, tiles or truncates it to 64\,600 samples (about 4\,s), scores it with the model under test and computes the equal error rate (\eer) from the resulting bona fide and spoof score lists. Our systems are wrapped as toolkit models, and the score written for each utterance is the logit difference $d = z_{\mathrm{bonafide}} - z_{\mathrm{spoof}}$, so that a higher score means bona fide, which is the convention the toolkit assumes. We use the toolkit's own \eer{} routine throughout and report \eer{} in percent; lower is better. Where a configuration was trained with three seeds we report the mean and the sample standard deviation over seeds and, where it matters, the per-seed range.

For each main Audio-JEPA and AudioMAE configuration, we evaluate the final six-epoch checkpoint without early stopping or checkpoint selection. A05 is monitored only; exploratory variants were scored on evaluation data and informed configuration choices. Both encoders use the same final downstream recipe.

\subsection{Datasets}
\label{sec:datasets}

Training uses the ASVspoof2019 LA training partition~\cite{todisco2019asvspoof,wang2020asvspoof}: 25\,380 utterances (2\,580 bona fide, 22\,800 spoofed) from 20 speakers and six attack algorithms, A01--A06: four text-to-speech systems and two voice conversion systems. With 20 speakers and six attacks, a random utterance split leaves both sides sharing speakers, attacks and source recordings, so we hold out an attack instead. We hold out A05 --- voice conversion with a WORLD vocoder --- together with 600 bona fide utterances, leaving 20\,980 utterances for training. Its vocoder is shared with A02 and A03 and its family with A06; A05 is used only for monitoring, and excluding it changes the training partition. The bona fide speech of ASVspoof2019 is drawn from the VCTK corpus~\cite{veaux2017vctk}.

Evaluation uses the five corpora in Table~\ref{tab:datasets}, all scored through the toolkit's protocol files. They differ in the origin of their genuine speech as well as in the kind of spoofing, and both distinctions matter for reading the results.

\begin{table}[t]
\centering
\caption{Evaluation corpora, and the AASIST baseline re-run through our pipeline against the figure published in Table~2 of~\cite{dowerah2025arena} (\eer{} in \%).}
\label{tab:datasets}
\footnotesize
\begin{tabular}{@{}lrl>{\raggedright\arraybackslash}p{2.7cm}rr@{}}
\toprule
 & & & & \multicolumn{2}{c}{AASIST} \\
\cmidrule(l){5-6}
Corpus & Utterances & Genuine & Spoofed speech & Ours & Publ. \\
\midrule
ASVspoof2019 LA eval~\cite{todisco2019asvspoof} & 71\,237 & VCTK & 13 TTS/VC attacks (11 unseen) & 0.83 & 0.82 \\
DFADD~\cite{du2024dfadd} & 3\,755 & VCTK & 5 diffusion / flow TTS systems & 39.05 & 41.86 \\
ADD 2022 Track~1~\cite{yi2022add} & 109\,199 & Mandarin & noisy, low-quality fakes & 47.92 & 47.91 \\
LibriSeVoc~\cite{sun2023librisevoc} & 18\,487 & LibriTTS & 6 neural vocoders & 37.95 & 37.65 \\
In-the-Wild~\cite{muller2022itw} & 31\,779 & web sources & fakes collected from the internet & 43.02 & 43.00 \\
\bottomrule
\end{tabular}
\end{table}

\emph{ASVspoof2019 LA evaluation partition} (7\,355 bona fide, 63\,882 spoofed) contains thirteen attacks, A07--A19; two of them (A16, A19) reuse the algorithms of training attacks A04 and A06, the other eleven are unseen. \emph{DFADD} (755 bona fide, 3\,000 spoofed) pairs VCTK speech with five recent diffusion- and flow-matching-based TTS systems, 600 utterances each; we use the copy distributed through HuggingFace,\footnote{\url{https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/DFADD}} a point we return to in Sect.~\ref{sec:rendition}. \emph{ADD 2022 Track~1} (31\,334 bona fide, 77\,865 spoofed) is Mandarin, low quality, and contains real-world noise and background music. \emph{LibriSeVoc} re-synthesises LibriTTS~\cite{zen2019libritts} utterances with six neural vocoders, so its spoofed speech is clean vocoder output on read speech. \emph{In-the-Wild} (19\,963 bona fide, 11\,816 spoofed) collects genuine and fake recordings of public figures from the internet.

%% VN: các số bona fide/spoof của In-the-Wild lấy từ Müller et al. (19 963 / 11 816 = 31 779 ✓).
%%     LibriSeVoc chưa có phân rã thật/giả trong tài liệu, nên không ghi.

\subsection{Baseline re-run and pipeline verification}
\label{sec:verification}

Before reporting our own numbers we checked the pipeline in three ways.

First, the leaderboard releases per-utterance score files for its systems. Recomputing \eer{} from those files with our own label files reproduces Table~2 of~\cite{dowerah2025arena} to within 0.01 point for seven of the eight released score sets we checked on ADD 2022 Track~1; this checks the metric and our label mapping without touching any audio.

Second, we ran the official AASIST checkpoint~\cite{jung2022aasist} through our installation of the toolkit on all five corpora (last two columns of Table~\ref{tab:datasets}), and AASIST is re-run alongside every one of our own evaluations. On ASVspoof2019 LA, ADD 2022 Track~1 and In-the-Wild the re-run matches the published figure to within 0.02 point; on LibriSeVoc it is within 0.30 of the figure in the paper and within 0.07 of the current online leaderboard value (38.02). On DFADD it does not match, with a known difference in audio rendition (Sect.~\ref{sec:rendition}).

Third, for each checkpoint we score the same waveform once through the toolkit wrapper and once directly in the training code and require the two scores to agree to $10^{-3}$, which checks that the wrapper reproduces the trained model.


\subsection{The DFADD audio rendition}
\label{sec:rendition}

On DFADD our AASIST re-run gives 39.05\pct{} against the published 41.86\pct. The leaderboard's released DFADD score file matches our copy on an identical list of 3\,755 files, and recomputing \eer{} from that file with our labels --- recoverable from the DFADD file names --- gives 41.83\pct, reproducing the published figure: same files, same labels. One known difference is the rendition of the audio, the leaderboard scoring a copy the organisers converted themselves (\texttt{test\_converted2} in their file paths) where we use the copy distributed through HuggingFace.

Where we place our DFADD result among published systems (Table~\ref{tab:dfadd}) we therefore use the ratio to the AASIST \eer{} measured in the same pipeline on the same audio --- 41.86 for the published systems, 39.05 for ours --- which cancels the rendition effect if it is multiplicative and common across systems. We have not tested that assumption, and we keep the raw \eer{}s in the table alongside.

\subsection{Comparison conditions}
\label{sec:conditions}

Leaderboard positions are only comparable between systems trained on the same data. Eleven of the twelve open-source systems reported in Table~2 of the Speech DF Arena paper~\cite{dowerah2025arena} are trained on ASVspoof2019 LA, as is ours; the exception is Whisper-MesoNet, which the arena paper lists separately. Entries added to the online leaderboard since then either do not disclose their training data or draw on a far broader corpus that includes the training splits of the evaluation sets themselves: Teffic-Audio~\cite{lin2026teffic}, for example, lists DFADD, LibriSeVoc, CodecFake and ADD 2022 among its training sources, and six leaderboard entries report 0.000\pct{} \eer{} on DFADD in the snapshot accessed on 9 September 2026. We therefore take baseline figures from Table~2 of the arena paper and compare only against those systems and against our own AASIST re-run. The online leaderboard~\cite{arenaleaderboard2026} is cited for context only, with its access date, since its contents change.

% ======================================================================
\section{Results}
\label{sec:results}

\subsection{Five benchmarks}
\label{sec:five}

Table~\ref{tab:main} gives \eer{} on the five corpora for the frozen and the fine-tuned Audio-JEPA system alongside the AASIST re-run.

\begin{table}[t]
\centering
\caption{\eer{} (\%) on five corpora. Audio-JEPA figures are mean $\pm$ sample standard deviation over three seeds unless marked; the frozen system trains only the 0.50\,M-parameter back-end, the fine-tuned system updates all 85.4\,M encoder parameters as well. AASIST is the official checkpoint re-run through the same pipeline.}
\label{tab:main}
\footnotesize
\begin{tabular}{@{}llrrr@{}}
\toprule
Corpus & Genuine speech & JEPA frozen & JEPA fine-tuned & AASIST \\
       &                & (0.5\,M trainable) & (85.9\,M trainable) & (0.3\,M) \\
\midrule
DFADD                & VCTK              & \textbf{8.47} $\pm$ 0.35 & 10.75 $\pm$ 3.48 & 39.05 \\
ASVspoof2019 LA eval & VCTK              & 17.70 $\pm$ 1.44 & \textbf{7.39} $\pm$ 0.69 & 0.83 \\
ADD 2022 Track~1     & Mandarin          & --- & \textbf{44.10} $\pm$ 2.04 & 47.92 \\
LibriSeVoc           & LibriTTS          & 47.38 $\pm$ 4.19 & 41.96 $\pm$ 2.93 & \textbf{37.95} \\
In-the-Wild          & web sources       & 47.55 $\pm$ 8.55 & 52.00$^{\dagger}$ & \textbf{43.02} \\
\bottomrule
\multicolumn{5}{@{}l@{}}{\footnotesize $^{\dagger}$one seed. --- not run.}
\end{tabular}
\end{table}

The clearest result is on DFADD. With the encoder frozen and only the back-end trained, the system reaches 8.47\pct{} (seeds 8.74, 8.59, 8.07) against 39.05\pct{} for AASIST. Fine-tuning the encoder does not help here: it gives 10.75\pct{} with ten times the seed spread (7.57, 10.20, 14.47).

On the ASVspoof2019 LA evaluation partition the picture reverses. Fine-tuning brings the system from 17.70 to 7.39\pct, against 0.83\pct{} for AASIST. This is an in-corpus evaluation, on a partition disjoint from training.

On ADD 2022 Track~1 the fine-tuned system is 3.8 points better than AASIST (44.10 against 47.92), in a corpus where the ASVspoof2019-trained systems in the arena paper span 31.04 to 50.35\pct. On LibriSeVoc and In-the-Wild the system is behind AASIST in both regimes; on In-the-Wild the frozen seeds range from 37.74 to 53.39, so a single seed there would say almost anything.

The two corpora on which the system reaches its lowest \eer{}s, DFADD and ASVspoof2019, are the two whose genuine speech comes from VCTK, the same corpus as the training data; the three corpora whose genuine speech comes from elsewhere --- LibriTTS, the internet, Mandarin speakers --- are where it is weakest. Sect.~\ref{sec:discussion} returns to this. 

Table~\ref{tab:dfadd} places the DFADD result among the twelve open-source systems of the arena paper. For the reason given in Sect.~\ref{sec:rendition}, the comparison column is the ratio of the AASIST \eer{} to each system's \eer{}, both measured in the same pipeline on the same rendition of the audio.

\begin{table}[t]
\centering
\caption{DFADD: systems trained on ASVspoof2019 LA, from Table~2 of~\cite{dowerah2025arena}, and ours. The ratio column divides the AASIST \eer{} measured in the same pipeline (41.86 for published systems, 39.05 for ours) by the system's \eer; higher is better and the rendition effect cancels under the assumption of Sect.~\ref{sec:rendition}. Parameter counts as reported by the arena paper.}
\label{tab:dfadd}
\footnotesize
\begin{tabular}{@{}rlrrr@{}}
\toprule
\# & System & Params & \eer{} (\%) & AASIST / \eer \\
\midrule
1 & XLSR+SLS~\cite{zhang2024xlsrsls}          & 340\,M & 7.54  & 5.55$\times$ \\
2 & TCM~\cite{truong2024tcm}                   & 319\,M & 8.88  & 4.71$\times$ \\
\textbf{3} & \textbf{Audio-JEPA frozen (ours)}   & \textbf{85.9\,M}$^{\ast}$ & \textbf{8.47} & \textbf{4.61$\times$} \\
4 & XLSR-Mamba~\cite{xiao2025xlsrmamba}        & 319\,M & 10.69 & 3.92$\times$ \\
5 & Nes2Net~\cite{liu2025nes2net}              & 318\,M & 11.14 & 3.76$\times$ \\
6 & Audio-JEPA fine-tuned (ours)               & 85.9\,M & 10.75 & 3.63$\times$ \\
7 & wav2vec2-AASIST~\cite{tak2022w2v2aasist}   & 318\,M & 11.92 & 3.51$\times$ \\
8 & RawGAT-ST~\cite{tak2021rawgatst}           & 0.44\,M & 23.70 & 1.77$\times$ \\
9 & Whisper-MesoNet~\cite{kawa2023whisper}$^{\ddagger}$ & 7.6\,M & 24.11 & 1.74$\times$ \\
10 & RawNet2~\cite{tak2021rawnet2}             & 17.6\,M & 26.19 & 1.60$\times$ \\
11 & WavLM-ECAPA~\cite{chen2022wavlm,desplanques2020ecapa} & 102\,M & 29.53 & 1.42$\times$ \\
12 & HuBERT-ECAPA~\cite{hsu2021hubert,desplanques2020ecapa} & 102\,M & 34.56 & 1.21$\times$ \\
13 & AASIST~\cite{jung2022aasist}              & 0.30\,M & 41.86 / 39.05 & 1.00$\times$ \\
14 & wav2vec2-ECAPA~\cite{baevski2020wav2vec2,desplanques2020ecapa} & 324\,M & 75.06 & 0.56$\times$ \\
\bottomrule
\multicolumn{5}{@{}l@{}}{\footnotesize $^{\ast}$0.5\,M of them trained. $^{\ddagger}$Not trained on ASVspoof2019 according to~\cite{dowerah2025arena}.}
\end{tabular}
\end{table}

\subsection{Pre-training objective: latent prediction against masked reconstruction}
\label{sec:objective}

Table~\ref{tab:objective} gives the four comparisons we ran.

\begin{table}[t]
\centering
\caption{Comparison of two released checkpoints under an identical downstream recipe. \eer{} (\%), mean $\pm$ s.d. over three seeds per arm, with per-seed ranges. Gap is AudioMAE minus Audio-JEPA; positive favours latent prediction.}
\label{tab:objective}
\footnotesize
\begin{tabular}{@{}llrrrrr@{}}
\toprule
Corpus & Encoder & JEPA & MAE & Gap & JEPA range & MAE range \\
\midrule
ASV2019 LA & fine-t. & 7.39$\pm$0.69  & 13.25$\pm$0.57 & $+5.87$ & 6.78--8.13 & 12.61--13.70 \\
ASV2019 LA & frozen     & 17.70$\pm$1.44 & 25.25$\pm$0.68 & $+7.55$ & 16.18--19.04 & 24.47--25.74 \\
DFADD           & frozen     & 8.47$\pm$0.35  & 11.04$\pm$0.75 & $+2.57$ & 8.07--8.74 & 10.18--11.54 \\
ADD22 Track~1   & fine-t. & 44.10$\pm$2.04 & 42.02$\pm$2.85 & $-2.08$ & 41.75--45.29 & 39.00--44.65 \\
\bottomrule
\end{tabular}
\end{table}

Audio-JEPA is ahead in three of the four comparisons, by 2.57 to 7.55 points, and in each of the three the seed ranges of the two encoders do not overlap: two corpora, two adaptation regimes, the same direction. In the fourth comparison, ADD 2022 Track~1 with fine-tuning, AudioMAE has the lower mean, by 2.08 points, with seed ranges that overlap over most of their extent.

On ADD 2022 Track~1 the first seed alone puts Audio-JEPA ahead by 0.67 points, where the three-seed means put AudioMAE ahead by 2.08. We therefore report non-overlapping seed ranges, rather than differences between means, as our descriptive criterion; three comparisons meet it and one does not. ADD 2022 Track~1 is also the noisy, low-quality corpus of the four, and the three comparisons that separate the encoders are on clean speech. Sect.~\ref{sec:discussion} offers an account of why that might be.

\section{Discussion}
\label{sec:discussion}

\paragraph{One account of the difference between the checkpoints.} Masked autoencoding minimises a reconstruction error in the input space. Under a squared loss the optimal prediction for a masked patch is its conditional expectation given the context, which regresses towards the conditional mean: components predictable from their surroundings survive, and those with high conditional variance are averaged away. The encoder only has to carry what that prediction needs. One hypothesis is that some discriminative spectral detail is poorly predictable from context; neither this property nor its retention by the encoders was measured here. A latent-predictive objective predicts in representation space, with targets the model itself produces; nothing there requires the representation to be sufficient for reconstruction, and nothing requires it to discard fine texture either.

This is a hypothesis about why two checkpoints differ, not a demonstration: the checkpoints differ in more than their objective (Sect.~\ref{sec:matched}), and we can only report whether the pattern it predicts appears. It does, in the direction predicted. The gap should be largest where the discriminative cue is fine spectral texture in clean audio and should close where noise masks that texture, and Table~\ref{tab:objective} shows 5.87 and 7.55 points on the clean ASVspoof2019 partition, 2.57 on DFADD, and no separation on the noisy ADD 2022 Track~1.

\paragraph{Association with the source corpus.} DFADD and ASVspoof2019 draw their bona fide speech from VCTK, as does the training data, and these are the two lowest \eer{}s; LibriTTS, web sources and Mandarin recordings do not, and these are the three highest.

Two things follow. Sharing downstream data does not eliminate encoder--corpus interactions. The VCTK association and the gap between checkpoints are observations, not a causal decomposition. And matched conditions are not by themselves enough to explain DFADD: AASIST is also trained on VCTK bona fide speech and reaches only 39.05\,\% there. LibriSeVoc, clean vocoder output on a corpus we did not train on, is where the two pull in opposite directions, and we did not run the AudioMAE arm there to see which wins.

\paragraph{What this implies for the pre-training corpus.} The front-ends above us on the leaderboard are speech models pre-trained on tens to hundreds of thousands of hours of speech. Audio-JEPA reaches the results above from 5\,338 hours of general audio --- between one and two orders of magnitude less --- which places the pre-training corpus among the levers this line of work has not used. The encoder's own evaluation points the same way: across the X-ARES suite its strengths lie in environmental sound and music rather than in speech~\cite{tuncay2025audiojepa}. We are not aware of a public speech-pretrained Audio-JEPA checkpoint. One of the strongest recent leaderboard entries, Teffic-Audio~\cite{lin2026teffic}, attributes its results to a training recipe of multi-source data, balanced sampling and heavy augmentation rather than to model design --- precisely the levers our single-corpus, lightly augmented recipe does not use.

% ======================================================================
% ======================================================================
\section{Future Work}
\label{sec:future}

Four directions follow from the results. First, the released checkpoint carries a trained predictor that we never use. Masking part of an utterance, predicting the representation of the masked part and scoring by prediction error could give a detector that needs no spoofed training data at all. Reconstruction error has long been used this way, but it is measured in the input space; a latent predictor scores in representation space instead, and the account in Sect.~\ref{sec:discussion} makes a testable prediction about which suits this task better. Second, the masking strategy is unexplored in the direction this task suggests: Audio-JEPA masks 40--60\,\% of patches at random, having found I-JEPA's block masking to perform worse on audio~\cite{tuncay2025audiojepa}, and A-JEPA~\cite{fei2023ajepa} anneals towards a SpecAugment-like scheme. Third, the effect of changing the pre-training corpus remains untested here. Fourth, training on genuine speech from more than one corpus remains untested here.

% ======================================================================
\section{Limitations}
\label{sec:limitations}

The comparison of Sect.~\ref{sec:matched} is between two released checkpoints, not between two objectives holding all else equal. The two share a backbone and a pre-training corpus and receive an identical downstream recipe, but their pre-training recipes were set independently and differ in masking ratio, sample rate, optimiser and compute budget, and AudioMAE receives 256-frame inputs rather than the 1\,024-frame inputs it was pre-trained on, which may cost it some accuracy. The account in Sect.~\ref{sec:discussion} attributing the difference to the objective is therefore a hypothesis consistent with the pattern we observe, not a controlled result. It rests on four measurements across three corpora, two clean and one noisy; we did not run the AudioMAE arm on LibriSeVoc or In-the-Wild.

The evaluation sets used to inform configuration choices were not held out from model development. The system is trained on a single corpus, and three of the five evaluation corpora draw their genuine speech from elsewhere; the results on those three characterise that dependence rather than the ceiling of the approach. The encoder has 85.4\,M parameters against 318--340\,M for the leading systems, so we have not tested latent prediction at a competitive scale. Our back-end differs from those of the published systems, so comparisons with them are not controlled. The DFADD figure is a system-level result and does not establish that the pretrained representation is necessary for that performance. Three seeds per configuration support the descriptive statements we make; we report no significance tests. We used no RawBoost or codec augmentation, which the field routinely applies. Finally, the DFADD placement in Table~\ref{tab:dfadd} rests on the ratio assumption of Sect.~\ref{sec:rendition}, and on one corpus.

% ======================================================================
\begin{credits}
%% VN: tác giả xác nhận không có tài trợ, nên bỏ hẳn \ackname. Disclosure là mục
%%     riêng và llncs v2.25 vẫn yêu cầu, nên giữ.
\subsubsection{\discintname}
The authors have no competing interests to declare that are relevant to the content of this article.
\end{credits}

% ======================================================================
\bibliographystyle{splncs04}
\bibliography{refs}

\end{document}
