PAGE 3 predicts cluster identities assigned offline by𝑘-means.Masked reconstruction(Au- dioMAE[13])decodesmaskedspectrogrampatchesbacktotheirvaluesandminimises theerrorininputspace.Latentprediction(I-JEPA,Audio-JEPA)usescontinuoustargets supplied by an exponential moving average of the encoder: with a context encoder𝑓𝜃, apredictor𝑔 𝜙 andatargetencoder𝑓 ¯𝜃 updatedasanexponentialmovingaverageof𝑓 𝜃, it minimises L= 1 |𝑀| 𝑔 𝜙 (︁𝑓 𝜃 (𝑥ctx), 𝑚)︁−sg [︁ 𝑓 ¯𝜃 (𝑥) ]︁ 𝑀 2 2,(1) where𝑚encodes the positions of the masked target patches andsgstops gradients. The target encoder processes the full input𝑥; its outputs at the masked positions𝑀 provide the targets. Latent prediction may allow representations to omit unpredictable inputdetails;whetherthisoccursandhelpsorhurtssynthesis-artefactdetectionremains an empirical question. Speech DF Arena.Dowerah et al. [8] released an evaluation toolkit and a leaderboard covering fourteen deepfake corpora. Their paper reports twelve open-source systems, all trained on ASVspoof2019 LA (Whisper-MesoNet [15] excepted), with released per- utterance score files; this is what makes an independent re-run checkable. We evaluate on five of the fourteen corpora (Sect. 4.2), chosen to span genuine-speech sources and spoofing types. 3 Method 3.1 Front-end: Audio-JEPA Audio-JEPA [30] applies the joint-embedding predictive architecture of I-JEPA [1] to log-mel spectrograms. During pre-training a context encoder sees a subset of spectro- gram patches, and a small predictor is trained to output the representations that a target encoder — an exponential moving average of the context encoder — assigns to masked patches. The loss is computed in representation space; no patch is ever reconstructed. The published checkpoint3 follows the ViT-Base configuration [7] (12 blocks, width 768,12heads,85.4Mparameters)andispre-trainedonAudioSet-2M[11]:5338hours of general audio spanning speech, music and environmental sound [30]. We use the context encoder and do not use the target encoder or the predictor.4 Thetoolkithandsthemodela16kHzwaveformtiledortruncatedto64600samples (Sect. 4.1). We resample to 32kHz, following the Audio-JEPA convention — this adds no information, since the source is 16kHz — take the first 2.56s (tiling if shorter), and compute a 128-band log-mel spectrogram with a 25ms window, 10ms hop and a 20–8000Hz range, after subtracting the waveform mean; the spectrogram is zero- padded or truncated to 256 frames. The256×128spectrogram is cut into patches of 8 frames by 32 mel bins (eight 10ms frame steps×a quarter of the mel axis), giving 3 https://huggingface.co/ltuncay/Audio-JEPA 4 BothI-JEPAandAudio-JEPAevaluatewiththetargetencoder,whichisanexponentialmoving average of the context encoder. We use Audio-JEPA’s context encoder; the effect of choosing it rather than its target encoder was not measured. PAGE 4 a grid of32×4=128tokens, each projected linearly to 768 dimensions and given a two-dimensional sinusoidal position embedding; there is no class token. The pre- trainedpatchembeddingis16×16;weinterpolateitsweightsbicubicallyto8×32and regenerate the position embeddings for the new grid. The token count is unchanged by this:Audio-JEPA’s16×16patchesovera128×256spectrogramgive128patches,and our(8,32)patches over256×128give32×4, also 128. These counts refer to the full patch grids before masking. These input settings differ from the pre-training recipe, which uses 10s clips at 32kHz with 128 mel bands over 256 frames [30] — about a 39ms hop, four times coarser in time than ours, with the released configuration setting the upper frequency to16kHz. 5 Thereleasedconfigurationoperatesataten-secondscale.Inanexploratory comparison, it reduced EER on the held-out attack subset by about a factor of five but increased evaluation-partition EER; these observations informed our decision to retain the finer-resolution configuration. 3.2 Back-end Theback-endisdeliberatelyconventional.Weaggregatetheencoder’slayeroutputsus- inglearnedweightsandsummarisetheresultingtokensequencewithattentivestatistics pooling [21]. Both arms of Sect. 3.4 use this back-end unchanged. The encoder yields thirteen token sequences of shape128×768: the output of each of the twelve blocks and the final layer-normalised output. Three steps turn them into a score.Alearnedsoftmaxoverthirteenscalarsweightsthesequencesandsumsthem(13 parameters, initialised to zero and hence uniform after the softmax). Attentive statistics pooling [21] — a two-layer attention network (768→128→1, tanh) with a softmax over the 128 tokens — produces attention-weighted mean and standard deviation vec- tors,concatenatedto1536dimensions(98561parameters).AclassifierofLayerNorm, Linear1 536→256,GELU,Dropout0.2andLinear256→2givestwologits(397058 parameters). The back-end totals 495632 parameters, about 0.5M, or 0.58% of the en- coder.Thescorewrittenforeachutteranceisthelogitdifference𝑑=𝑧 bonafide −𝑧 spoof;we usethedifferenceratherthanasinglelogitbecausethesoftmaxprobabilityismonotone in the difference, not in either logit alone, and the toolkit’s EER depends only on score order. Figure 1 distinguishes the published pre-training procedure from the downstream pipeline used here. 3.3 Training Wetrainintworegimes.Frozen:theencoderrunsinevaluationmodewithoutgradients and only the back-end is trained.Fine-tuned: both encoder and back-end are fine-tuned, withapeaklearningrateof3×10 −5 fortheencoder.Theback-endusesapeaklearning rateof10 −3 inbothregimes.BothuseAdamW[19]withweightdecay0.05,aone-cycle schedule with 10% warm-up, batch size 32, mixed precision and exactly six epochs. Thelossiscross-entropywiththebonafideclassup-weightedbythespoof-to-bona-fide 5 https://github.com/LudovicTuncay/Audio-JEPA PAGE 5 (a) Published Audio-JEPA pre-training Spectrogram Context encoder Predictor Target encoder Select masked positions Latent prediction loss Visible patches Full input EMA update Stop-gradient Target encoder sees the full input; targets are selected at masked positions. (b) Downstream detection in this work Waveform to tokens 16 to 32 kHz; 2.56 s 256 time frames × 128 mel bins Patch: 8 × 32 128 tokens × 768 Audio-JEPA context encoder Frozen: fixed; fine-tuned: trained Block 1 Blocks 2-11 Block 12 Final LN 13 readouts, each 128 tokens × 768 Back-end: trained in both regimes Layer aggregation Softmax-weighted sum 128 tokens × 768 Attentive stats pooling Mean + standard deviation 1,536 dimensions Classifier (two logits) 1,536 → 256 → 2 d = z bonafide - z spoof Fig.1.(a)Audio-JEPApre-training,schematicallyredrawnfrom[30]:thecontextbranchpredicts representationsatmaskedpositions;thetargetencoderseesthefullinputandisupdatedbyexpo- nentialmovingaverage(EMA).Gradientsdonotpassthroughthetargetbranch.(b)Downstream detection using the released context encoder; we do not repeat pre-training. Its twelve block out- puts and final layer-normalised output are aggregated. Only the back-end is trained in the frozen regime; both encoder and back-end are trained in the fine-tuned regime. ratio of the training set (about 9.6). The only augmentation, applied during training, is a SpecAugment-style [22] mask drawn afresh for each spectrogram: one band of up to 16 of the 256 frames and one of up to 16 of the 128 mel bins, at random positions — at most 6.25% of the time axis and 12.5% of the mel axis. This differs from the patch masking used during encoder pre-training, which masks 40–60% of patches to define PAGE 6 thelatent-predictiontask.Trainingdataandhold-outaredescribedinSect.4.2;thefinal six-epoch checkpoint is used for each main configuration (Sect. 4.1). Every configuration in Tables 2 and 4 is trained with three seeds, 1234, 2 and 3, except where a single seed is marked. We report both the frozen and the fine-tuned regime, not to choose between them but because they give that comparison a second axis: a difference that survives with the encoder frozen cannot be an artefact of how fine-tuning interacts with one checkpoint or the other. 3.4 Comparison against a reconstruction-trained checkpoint TheclosestavailablepointofcomparisonforAudio-JEPAisAudioMAE[13]:thesame ViT-Base backbone, the same AudioSet-2M pre-training corpus, and a reconstruction objective in place of a predictive one. We take the encoder of the public checkpoint (thetimmportvit_base_patch16_1024_128.audiomae_as2m),interpolateitspatch embeddingto8×32inthesameway,andgiveitthesameinput,back-end,trainingdata, recipe and seeds. The checkpoints share the ViT-Base architecture and AudioSet-2M as their pre- training data source, and we use a common downstream recipe. Their pre-training recipesneverthelessdifferinmaskingratio,samplerate,optimiserandcomputebudget, among other factors. The two encoders are also displaced from their pre-training input configuration along different axes: our 2.56s input shortens the clip about fourfold for both, preserves Audio-JEPA’s full patch count while quartering AudioMAE’s, and changes the nominal temporal patch span (patch length in frames multiplied by frame hop)from625to80msforAudio-JEPAandfrom160to80msforAudioMAE.Nosingle configurationisnativetoboth,andgivingeachitsownwouldchangetheinputalongside theobjective.Wethereforereadthecomparisonasonebetweentworeleasedcheckpoints under a common downstream recipe, and treat the attribution of the difference to the objective itself as a hypothesis, examined in Sect. 6 and bounded in Sect. 8. 4 Experimental Setup 4.1 Evaluation toolkit and protocol AllevaluationsusetheSpeechDFArenatoolkit[8].Inthefixed-lengthconfigurationwe run, the toolkit loads every file at 16kHz, tiles or truncates it to 64600 samples (about 4s), scores it with the model under test and computes the equal error rate (EER) from theresultingbonafideandspoofscorelists.Oursystemsarewrappedastoolkitmodels, and the score written for each utterance is the logit difference𝑑=𝑧bonafide −𝑧 spoof, so that a higher score means bona fide, which is the convention the toolkit assumes. We usethetoolkit’sownEERroutinethroughoutandreportEERinpercent;lowerisbetter. Where a configuration was trained with three seeds we report the mean and the sample standard deviation over seeds and, where it matters, the per-seed range. ForeachmainAudio-JEPAandAudioMAEconfiguration,weevaluatethefinalsix- epoch checkpoint without early stopping or performance-based checkpoint selection. During these runs, the A05 hold-out is monitored but does not determine which check- point is evaluated. Exploratory variants were scored on evaluation data and informed configuration choices. Both encoders use the same final downstream recipe.