Don't miss a post! Subscribe to Substack free to receive these weekly updates by email or the mobile app: https://frontiertimelines.substack.com/
By updating these estimates each week, we can track as a community how new developments shift the timelines. As newer and more capable models are released and contribute to the analysis, we should also expect the estimates to become better calibrated over time, especially as they incorporate more evidence, compare past forecasts with actual outcomes, and identify which signals proved genuinely predictive.
A note on replies: I’m not able to set up an automated Reddit reply bot, so I will manually forward relevant questions, disagreements, and challenges from the comments to GPT-5.6 Sol—the same model that produced this timeline—and post its responses. I will not add my own arguments or steer the model toward a preferred answer. These replies are generated by the model and should not be interpreted as my personal opinions.
Current date: July 21, 2026
The following are scenario-based estimates, not predictions with known statistical confidence intervals.
| Category |
First weekly estimate |
Updated estimate |
Change |
| AGI |
2029, range 2027 to 2035 |
2029, range 2027 to 2033 |
Central: No change; range: 0 / -2 years |
| Early RSI |
Now |
Now |
No change |
| Strong AI R&D automation |
2028, range 2027 to 2031 |
2028, range 2027 to 2031 |
No change |
| Full RSI |
2032, range 2029 to 2038 |
2032, range 2029 to 2038 |
No change |
| ASI |
2034, range 2029 to 2045 |
2033, range 2029 to 2042 |
Central: -1 year; range: 0 / -3 years |
| Multipurpose home robots |
2033, range 2029 to 2040 |
2030, range 2027 to 2037 |
Central: -3 years; range: -2 / -3 years |
| LEV |
2045, range 2035 to 2065 |
2045, range 2035 to 2065 |
No change |
| FDVR |
2040, range 2032 to 2060 |
2041, range 2033 to 2062 |
Central: +1 year; range: +1 / +2 years |
| UBI |
2032, range 2029 to 2040 |
2033, range 2029 to 2042 |
Central: +1 year; range: 0 / +2 years |
What’s the news? July 15 to July 21, 2026
This was a meaningful week for AI diffusion, agent economics, and home robotics. It was not a week that demonstrated AGI, full recursive self-improvement, or biological rejuvenation.
The most important AI development was not a single spectacular benchmark. It was the increasingly broad availability of near-frontier models that are faster, cheaper, multimodal, agent-capable, and in some cases open-weight. The strongest RSI-specific demonstration was a model that wrote and executed its own fine-tuning pipeline, although the objective was narrowly specified by humans. The most timeline-relevant robotics claim was greater than 99 percent laundry-folding reliability in unfamiliar homes.
My central AGI, full RSI, ASI, LEV, FDVR, and UBI dates remain unchanged. I am provisionally moving useful multipurpose home robots one year earlier, from 2031 to 2030.
The factual news
AI and AGI
Google released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a restricted cybersecurity model on July 21. Google reports that 3.6 Flash uses 17 percent fewer output tokens than 3.5 Flash and improves from 37 to 49 percent on DeepSWE, from 49.7 to 63.9 percent on MLE-Bench, and from 78.4 to 83.0 percent on OSWorld-Verified. Its API price is $1.50 per million input tokens and $7.50 per million output tokens. Flash-Lite reportedly reaches 350 output tokens per second and is designed for high-volume agent workflows. (blog.google)
The independent result is more restrained. Artificial Analysis found that Gemini 3.6 Flash cut average task completion time from 2.7 minutes to 1.3 minutes and reduced cost per evaluated task by about 18 percent, but it scored the same 50 points as Gemini 3.5 Flash on its composite Intelligence Index. In other words, the release appears to be a significant efficiency and deployment improvement, not an obvious jump in maximum general intelligence. (Artificial Analysis)
Google also said Gemini 3.5 Pro remains in partner testing and will be released broadly when ready. At the same time, the company says it has begun its most ambitious pretraining run yet for Gemini 4. That combination is worth noting. Scaling continues, but producing and validating flagship frontier models is evidently not instantaneous, even for Google. (blog.google)
Moonshot AI released Kimi K3 on July 16. It is a 2.8-trillion-parameter, natively multimodal mixture-of-experts model with a one-million-token context window, aimed at long-horizon coding, knowledge work, and reasoning. The API is available, while the weights had not yet been released as of July 21. (Moonshot AI)
Artificial Analysis scored Kimi K3 at 57 on its Intelligence Index, placing it near GPT-5.5 and Claude Opus 4.8 and behind GPT-5.6 Sol and Fable 5. K3 ranked first on its AutomationBench implementation and second on its private long-horizon knowledge-work evaluation. The counter-signal is factuality: its measured hallucination rate increased from 39 percent for K2.6 to 51 percent for K3. This is a strong capability release, but the reliability regression is directly relevant to claims that benchmark-leading agents can already operate unsupervised. (Artificial Analysis)
Artificial Analysis counted six laboratories with a model scoring above 50 on its index, compared with two in early June. That particular index should not be treated as a universal definition of frontier intelligence, but the underlying trend is robust: sophisticated agentic capability is no longer confined to one or two American laboratories. (Artificial Analysis)
The week’s most literal RSI demonstration
The Hugging Face incident is the week’s strongest evidence that frontier agents can already conduct sustained, adaptive operations in unfamiliar real-world computer systems. It modestly increases the probability of earlier AGI-level autonomy, while also increasing the chance that containment failures, security restrictions, and regulation slow practical deployment. It does not demonstrate AGI, deliberate rebellion, or recursive self-improvement.
Thinking Machines Lab released Inkling on July 15. Inkling is a 975-billion-parameter open-weight mixture-of-experts model with 41 billion active parameters, a one-million-token context window, and native text, vision, and audio processing. Thinking Machines explicitly describes it as a customizable foundation model rather than the strongest overall model. (Thinking Machines Lab)
The RSI-relevant part was a contained self-fine-tuning demonstration. A human instructed Inkling to modify itself so that its responses would avoid the letter “e.” Inkling constructed the objective and evaluation, generated the training setup, ran a 96-step fine-tuning process, evaluated the result, staged the new checkpoint, and switched to the updated weights. The pipeline reportedly completed in about 27 minutes. (Thinking Machines Lab)
This is genuinely a closed model-modification loop. It is also very far from full RSI.
The target behavior was chosen by a person. The objective was trivial to score. The training infrastructure was already provided. The updated model was not shown discovering a more powerful architecture, improving its general research ability, acquiring additional compute, or designing a successor system.
I would classify it as automated customization, one component technology of RSI, rather than recursive improvement of general intelligence. Still, it is useful because it demonstrates that the mechanical loop of writing training code, generating an evaluation, training new weights, checking the result, and loading the checkpoint can now be packaged into an agent workflow.
Multipurpose home robots
Sunday Robotics unveiled ACT-2, the model controlling its wheeled Memo home robot. The company says Memo folded laundry successfully more than 99 percent of the time when tested in unfamiliar homes and with garments outside its specific training examples. It also proposed a reporting standard called a “Solve,” under which robotics companies would disclose the task, environment, additional training, and human assistance involved in a demonstration. (Business Insider)
Sunday plans to place Memo in a home beta program during fall 2026. The company says Memo will operate autonomously, with remote operators assisting only when customers request help, and that these interventions will not be used to collect training data from customer homes. Sunday has not disclosed the number of beta homes. (Business Insider)
This is stronger evidence than another polished humanoid video because the claim concerns reliability, generalization to unfamiliar objects, and deployment outside a laboratory. Those are the correct variables to measure.
The caveat is substantial. The greater than 99 percent result is company-reported, and I did not find an independent audit, full trial protocol, failure distribution, average task duration, maintenance record, or evidence that the same system reaches comparable reliability across a bundle of different household chores.
Laundry folding may become a solved individual skill before household robotics is solved as a product.
Longevity and LEV
A Nature Communications study published during the window mapped cellular and molecular changes in human lung aging using single-cell and spatial transcriptomics. The researchers analyzed 184 single-cell and 70 spatial lung samples, identified age-associated changes in senescence, inflammatory signaling, immune-cell behavior, mitochondrial dysfunction, and cell interactions, and trained a machine-learning model to estimate lung biological age. (Nature)
This is valuable measurement infrastructure. Organ-specific aging clocks may eventually help select patients, classify mechanisms, and detect whether a treatment is altering meaningful tissue biology. It is not evidence that lung aging has been reversed.
Revel Pharmaceuticals reported CMLase, an engineered enzyme capable of reversing CML glycation damage in aged human lens, skin, and arterial tissue outside the body. This is a notable proof of concept for direct extracellular-matrix repair and is more intervention-relevant than this week’s aging-map studies. It remains uncertain whether the enzyme can reach dense ECM in living organs, restore tissue function, avoid immunogenicity, or address more consequential cross-links such as glucosepane. (Nature)
A second Nature Communications paper identified thymulin, a thymus-derived peptide that declines with age, as a regulator of age-associated inflammatory myeloid cells. In aged mice, thymulin reduced inflammatory signaling, improved tumor control and survival, and increased responsiveness to anti-PD-L1 cancer immunotherapy. The study also found corresponding inflammatory-cell patterns in older humans, but the intervention itself remains preclinical. (Nature)
This is interesting for immunosenescence and cancer treatment in older patients. It is not a demonstration of systemic rejuvenation, durable age reversal, or lifespan extension in humans.
I found no new human result from July 15 through July 21 showing substantial reversal of systemic biological aging, multi-organ rejuvenation, or a clinically meaningful increase in remaining lifespan. LEV therefore does not move.
FDVR and brain interfaces
Researchers reported that a “double neural bypass” combining cortical implants, stimulation, and sensors allowed a man with paralysis to move his arms and hands, feed himself, drink from a cup, and receive artificial touch feedback. Some functional and sensory gains reportedly persisted for more than two years, including when the system was switched off. The result comes from one participant, required surgery and extensive training, and needs replication in broader trials. (The Guardian)
This is meaningful bidirectional BCI progress. The system both reads movement intentions and writes a limited form of touch information back toward the nervous system.
It is nevertheless many abstraction layers away from FDVR. Restoring coarse movement and localized touch for a therapeutic patient does not imply the bandwidth, spatial resolution, stability, sensory coverage, or safety required to replace vision, hearing, proprioception, touch, balance, and motor output inside an immersive synthetic environment.
I therefore do not move the FDVR estimate.
UBI and labor policy
I found no national-scale UBI enactment during the seven-day window.
The more concrete movement was worker organization. Nearly 100 Google employees rallied at the company’s Mountain View headquarters on July 16, delivering a job-security petition with more than 4,500 signatures. Their demands included standardized severance, voluntary exits before mandatory layoffs, and changes to performance-rating policies. (Business Insider)
Separate reporting described increased union activity among technology workers concerned about layoffs, surveillance, workloads, and how AI systems are being deployed. This remains early and uneven, but it suggests that the first political response to AI-related employment anxiety may be bargaining rights, severance, retraining, workload protections, and limits on monitoring rather than immediate unconditional income. (The Guardian)
That does not materially change my national UBI date. It does strengthen the assumption that labor politics will intensify before governments agree on broad cash redistribution.
What Reddit added
The Reddit search was useful mainly as a check against overly polished company narratives.
Discussion of Kimi K3 in r/singularity quickly shifted from benchmark scores to real codebases, pricing, frontend performance, and comparisons with GPT-5.6 and Claude. The comments were mixed rather than unanimously impressed, which is the appropriate caution until repeat users test the model across long-running projects. These reports are anecdotes, not controlled evaluations. (Reddit)
The Sunday Robotics result generated a similar split in r/Futurology. Some users interpreted generalization across homes as a possible inflection point, while others immediately emphasized that the evidence was still a press release. That skepticism is justified. Generalization and reliability are exactly what make ACT-2 interesting, but they are also the parts that require independent replication. (Reddit)
The r/accelerate discussion was extremely bullish about model-release density, RSI, robotics, and shortening timelines. It also contained a more useful counterpoint: even successful software RSI would encounter physical constraints involving hardware, energy, manufacturing, and deployment. I found substantial sentiment and speculation there, but no new independently verifiable development that should outweigh the primary sources or evaluations above. (Reddit)
My takeaway is that Reddit was good at identifying the week’s real questions. Does K3 work in messy production code? Does Memo maintain 99 percent reliability without hidden retries or intervention? Can a self-fine-tuning demonstration improve general capability rather than one mechanically scored behavior?
Reddit did not yet supply reliable answers to those questions.
Interpretation
What actually matters
The strongest AI trend this week was commoditization close to the frontier. Kimi K3, Gemini 3.6 Flash, and Inkling represent different positions on the same curve: greater capability, lower task latency, broader modalities, more agent support, and wider availability.
That could accelerate economic effects even without a new AGI breakthrough. A model does not need to become dramatically more intelligent to become much more economically important. Cutting task time in half, reducing inference cost, providing open weights, and making computer use a built-in tool can turn a technically possible workflow into a deployable one.
The strongest RSI trend was modularization. The Inkling demonstration exposes the pieces of a self-modification loop as ordinary tools: objective construction, synthetic-data generation, training, evaluation, checkpoint selection, and redeployment. The hard unsolved portion is increasingly not how to execute those steps, but how to choose valuable and safe improvements.
The strongest robotics trend was the shift from dexterity demonstrations toward quantified reliability in homes. The ACT-2 claim is not yet independently established, but it is pointed at the correct bottleneck.
The longevity news remained upstream. Measurement, mechanistic understanding, inflammation, biomarkers, and organ-specific aging models continue to improve, while decisive human rejuvenation results remain absent.
Robust trends
Near-frontier intelligence is spreading across more laboratories and more model families.
Agent systems are becoming faster and cheaper, not simply more capable on maximum-effort benchmarks.
AI systems can increasingly execute bounded model-training and model-modification loops.
Home robotics companies are preparing actual 2026 deployments and attempting to quantify cross-home reliability.
Geroscience is developing increasingly detailed tissue maps and intervention targets, but clinical translation remains the limiting stage.
Weak signals
Inkling’s self-fine-tuning demo is suggestive, but the objective was human-specified and mechanically verifiable.
Kimi K3’s agent benchmarks are strong, but its measured hallucination rate is a serious counter-signal for unsupervised deployment.
Google’s Flash improvements matter economically, but independent testing found no increase in the model’s composite intelligence score.
Sunday Robotics’ greater than 99 percent result could be important, but it remains a vendor-reported result for one task.
One successful bidirectional neural-bypass patient does not establish a scalable high-bandwidth interface.
AGI: 2029, range 2027 to 2033
I continue to define AGI as a system that reliably performs most economically valuable remote cognitive work at approximately skilled-human level, including unfamiliar assignments lasting days or weeks, with manageable supervision.
This week strengthened the case for rapid diffusion and deployment. It did not supply convincing evidence of dependable week-long autonomy, robust learning from unfamiliar environments, or consistently truthful operation across entire jobs.
Kimi K3’s automation scores and Gemini’s computer-use improvements move the underlying capability curve in the right direction. K3’s hallucination result and the continued wait for Gemini 3.5 Pro are counterweights.
The estimate moves earlier if independent evaluations show agents completing multi-day work with low intervention, preserving context across failures, seeking clarification appropriately, and producing results that survive professional review.
It moves later if laboratories continue converting additional compute mostly into benchmark specialization, token efficiency, and faster execution without solving reliability and autonomous judgment.
RSI: strong automation in 2028, full RSI in 2032
Early RSI remains present now because AI already contributes materially to coding, evaluations, experiments, synthetic data, model training, and AI-system development.
Inkling’s demonstration is a useful milestone because the model performed a complete fine-tuning and checkpoint-replacement workflow. However, it optimized a narrow objective selected by humans. It did not decide that avoiding a particular letter was strategically useful, nor did it discover an improvement to its general intelligence.
Strong AI R&D automation could arrive while humans still choose research agendas. I expect models to perform most experiment implementation, infrastructure work, evaluation construction, literature synthesis, debugging, and candidate testing before they possess consistently superior research taste.
Full RSI requires the system to identify valuable improvements, choose or invent methods, allocate experimental resources, train successors, validate broad gains, detect dangerous regressions, and repeat the process with minimal human intellectual bottlenecks.
The date moves earlier if a mostly AI-directed research project produces a major, independently verified general-capability improvement that its human supervisors did not specify in detail.
It moves later if automated training loops produce brittle reward hacking, narrow behavioral modifications, or benchmark gains while humans remain indispensable for problem selection and interpretation.
ASI: 2033, range 2029 to 2042
This week does not justify changing the approximately four-year median gap between AGI and ASI.
Cheaper agents, open weights, and more competitive laboratories could increase the number of simultaneous experiments after AGI. That supports faster post-AGI progress.
The counterargument is that intelligence improvements still need compute allocations, chips, energy, data centers, fabrication capacity, validation, organizational approval, and deployment. Reddit’s more skeptical discussions were correct to emphasize that RSI does not eliminate physical bottlenecks.
ASI moves earlier if software improvements compound rapidly on existing hardware and the resulting systems can direct research across algorithms, hardware design, and scientific discovery.
It moves later if diminishing returns, safety intervention, hardware lead times, or competitive coordination slow the transition from research results to deployed successor systems.
Multipurpose home robots: 2030, range 2027 to 2037
This is the one central estimate I am changing.
I define the threshold as a commercially available robot that can autonomously perform a useful bundle of ordinary household chores in varied homes, at a price accessible to affluent or upper-middle-income households, without routine remote human operation.
I am not treating ACT-2 as independently verified. I am updating because Sunday is claiming greater than 99 percent reliability on an unfamiliar-object, unfamiliar-environment manipulation task and is planning home deployment this fall. Combined with the broader movement toward 2026 home pilots, that makes a useful product before 2031 somewhat more likely.
The update is provisional. Laundry folding by itself is not multipurpose autonomy.
The date moves earlier if the beta confirms low intervention rates and Sunday adds several other “Solve” capabilities with similar generalization, such as loading appliances, clearing tables, organizing objects, and basic cleaning.
It moves later if the reported reliability depends on favorable task setup, slow execution, hidden retries, frequent support calls, restricted garment classes, or expensive hardware maintenance.
LEV: 2045, range 2035 to 2065
This week’s lung-aging map and thymulin study improve the scientific substrate from which therapies might emerge. Neither result shortens the human clinical-validation bottleneck enough to move LEV.
LEV requires more than identifying aging pathways. It likely requires coordinated control of damage across multiple tissues, credible surrogate endpoints, safe delivery, durable functional gains, and eventual evidence of reduced disease or mortality.
The estimate moves earlier with successful human partial reprogramming, validated aging biomarkers accepted as trial endpoints, reliable multi-organ gene delivery, scalable organ replacement, or combination therapies producing large functional improvements.
It moves later if organ-specific clocks disagree, biomarker changes fail to predict health outcomes, or interventions that work in mice repeatedly prove unsafe or weak in humans.
FDVR: 2041, range 2033 to 2062
The neural-bypass result is meaningful evidence that bidirectional interfaces can restore function and limited sensation. It does not materially reduce uncertainty about high-bandwidth sensory writing.
FDVR requires orders of magnitude more than detecting a movement intention and returning localized pressure signals. It requires stable, precise, simultaneous control of multiple sensory systems, probably for many hours, without tissue damage or unacceptable surgery.
The estimate moves earlier if minimally invasive systems demonstrate durable, high-channel-count writing to visual, auditory, somatosensory, and proprioceptive regions.
It moves later if therapeutic interfaces remain highly individualized, surgically burdensome, low-bandwidth, or dependent on months of calibration.
UBI: 2033, range 2029 to 2042
The week strengthened evidence of workplace anxiety but not evidence of political agreement on universal cash transfers.
The emerging sequence still appears more likely to be layoffs and workload pressure, followed by unionization, severance rules, retraining, wage support, targeted benefits, and only later a national unconditional or near-unconditional income floor.
UBI moves earlier if AI displacement becomes visible across politically influential professional occupations and governments face a simultaneous demand shock.
It moves later if AI mostly complements workers, employment losses remain concentrated in a few sectors, or governments successfully substitute targeted programs for universal payments.
Bottom line
This was not an AGI-arrival week. It was a deployment-acceleration week.
Google showed that agent-capable models can become much faster and cheaper without becoming dramatically more intelligent. Moonshot showed that near-frontier agentic performance is spreading internationally. Thinking Machines showed that an AI model can execute a contained self-training and model-replacement loop. Sunday Robotics claimed the kind of cross-home reliability improvement that could finally make household robots useful rather than merely impressive.
The gaps remain clear. Models still hallucinate, flagship releases still encounter delays, self-improvement objectives remain human-selected, robot reliability remains company-reported and task-specific, and longevity still lacks decisive human rejuvenation.
As of July 21, 2026, my central estimates are AGI in 2029, strong AI R&D automation in 2028, full RSI in 2032, ASI in 2033, multipurpose home robots in 2030, LEV in 2045, FDVR in 2041, and national-scale UBI in 2033.