AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research describes Dust, a zeroth-order method that trains transformer language models by perturbing activations rather than calculating backpropagation gradients. The report says Dust approaches or sometimes exceeds backpropagation in tested settings, but larger populations require more compute and the findings do not establish a practical replacement for backprop at production scale.

Q Labs Research has introduced Dust, a method for pretraining transformer language models without backpropagation, the standard process for calculating training gradients. In a report dated October 2026, the researchers say Dust’s estimates become more similar to backpropagation as the method’s population grows and that it performs competitively in some tested settings, while acknowledging that this takes substantially more compute.

Dust uses zeroth-order optimization: it perturbs a model’s activations, measures how those perturbations affect the loss, and combines the perturbations according to their results to estimate an update. Unlike weight-space evolution strategies, it does not need to create and run a separate copy of the model for each candidate. Instead, the researchers say, perturbations are applied independently across tokens, allowing one forward pass to evaluate a virtual population in parallel.

The report says Dust’s estimates align more closely with backpropagation as population size increases, and remain well aligned across the scales tested, up to 1 billion tokens. It also reports that a 243-million-parameter model outperformed a model 120 times smaller at most tested population sizes. Those are findings from the researchers’ experiments, not evidence that the method is faster or better in every training setting.

Q Labs also compares Dust with EGGROLL, a weight-space evolution-strategy method. Based on extrapolations, the researchers estimate that from one million tokens upward, Dust is roughly 1,000 to 10,000 times more compute-efficient than a transformer implementation of EGGROLL. That comparison is against EGGROLL, not backpropagation, and the report characterizes it as an extrapolation rather than a direct measurement across all scales.

At a glance
reportWhen: Research report dated October 2026
The developmentQ Labs Research published a report on Dust, an activation-perturbation method for pretraining transformers without backpropagation.

A Different Route to Transformer Training

Backpropagation has shaped the architectures, optimizers and hardware used to train modern neural networks. A method that could train large models without a backward pass might broaden the optimization approaches researchers can test and could matter in settings where large amounts of compute are available. Dust’s token-level virtual population is intended to reduce a major cost of earlier search-based methods: evaluating many separate model variants.

The report does not show that Dust currently beats backpropagation on practical training cost or language-model quality. Instead, its importance is as an experimental result: activation perturbations can produce useful training signals for transformer models in the tested settings. Whether that advantage can be made reliable and economical at larger scales remains unresolved.

Amazon

AI training hardware accelerators

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Weight Search to Activations

Evolution strategies and other zeroth-order approaches search for useful parameter changes using evaluations of model performance rather than analytic gradients. Their drawback is that large populations can require many separate evaluations. Dust builds on node perturbation, applying random changes to intermediate activations instead of directly perturbing model weights.

The report frames this as a way to use more brute-force computation and less analytic structure than backpropagation. Its central proposal is that each token can act as a member of a virtual population, with the evaluations carried out together in a forward pass. The paper’s argument that this could become useful in compute-rich settings is a research hypothesis, not a demonstrated outcome for deployed language models.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, in the report’s abstract

Amazon

transformer model training GPUs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Tests

The report’s supplied summary does not establish how Dust compares with backpropagation on end-to-end training time, energy use, or final language-model quality across a broad set of model sizes and tasks. It also does not establish whether the method’s compute demands remain manageable as population size rises. The claimed advantage over EGGROLL is based on extrapolations, and the comparison basis should not be extended to other methods.

It remains unclear whether the results will be independently replicated, how sensitive they are to implementation and tuning choices, and whether Dust can scale beyond the reported tests while remaining competitive in total cost. The authors’ suggestion that larger models are more population-efficient describes their observed experiments, not a settled rule for future models.

Amazon

neural network training optimization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Replication and Larger-Scale Tests

The next evidence needed is further testing of Dust against backpropagation and other training methods under matched compute budgets, with model quality, wall-clock time and hardware costs reported together. Independent replication would help clarify whether the reported gradient alignment and model-size results hold beyond the authors’ experiments.

The report presents a research method, not a deployment announcement, and does not give a confirmed timeline for follow-up work. Until broader results are available, Dust is best understood as a candidate approach to transformer training whose practical value remains to be established.

Amazon

machine learning compute resources

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Dust?

Dust is a zeroth-order training method from Q Labs Research. It perturbs neural-network activations and uses the resulting changes in loss to estimate model updates, rather than calculating gradients with backpropagation.

Does Dust eliminate backpropagation?

The method is designed to train without a backward pass, but the report does not show that it has replaced backpropagation in practical or production training. Its reported results are experimental.

Did Dust outperform backpropagation?

Q Labs says Dust approached backpropagation’s estimates at larger population sizes and exceeded it in multiple tested settings. The summary does not establish a universal advantage in model quality, runtime or total compute cost.

How does Dust compare with EGGROLL?

The researchers estimate that Dust is about 1,000 to 10,000 times more compute-efficient than a transformer implementation of EGGROLL from one million tokens upward. They describe this as an extrapolation; it is not a comparison with backpropagation.

What remains to be shown?

Further evidence is needed on performance under matched compute budgets, total training cost, final model quality, scalability and independent replication. The report does not establish how Dust would perform across a wider range of models and tasks.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Scorching Heat Makes ‘Ipchu’ Meaningless… Seoul Tops 40 Degrees For The First Time In 8 Years – 경향신문

Seoul experiences its first 40°C temperature in eight years, rendering traditional ‘Ipchu’ weather predictions meaningless amid ongoing heatwave.

Will It Rain In Oklahoma City On Aug 26, 2026?

A prediction market indicates possible rain in Oklahoma City on August 26, 2026, but weather forecasts remain highly uncertain for that date.

Ten Advances In Mathematics And Theoretical Computer Science

A review of ten recent significant advances in mathematics and theoretical computer science, highlighting confirmed developments and their implications.

erdbeben neapel

A significant earthquake struck Naples today, causing structural damage but no confirmed casualties. Authorities are assessing the situation.