AI Plugin from Sakana AI Simplifies Document Processing for Large Models
Large language models' struggle with processing long texts, often called "memory anxiety," may soon be a thing of the past. Recently, Tokyo-based AI startup Sakana AI unveiled two breakthrough technologies: Text-to-LoRA (T2L) and Doc-to-LoRA (D2L). Utilizing an innovative "super network" architecture, these technologies enable large models to "absorb" massive documents or learn new tasks in under a second, all without retraining.

AI developers have long faced a difficult trade-off: cramming lengthy documents into a chat window—which slows responses and consumes high memory—or paying the high cost of fine-tuning a model. Sakana AI introduces a third option. Through a one-time payment for pre-training, it generates minuscule weight plugins (LoRA) to achieve low-cost, efficient model adaptation.
Doc-to-LoRA: Memory Needs Drop from 12GB to Just 50MB
This represents the most impressive technology in the latest release. Processing a 128,000-token document (roughly 100,000 words) with traditional methods demands over 12GB of VRAM to store information. With D2L, the model can directly "digest" this information into a plugin smaller than 50MB.
Remarkable Speed: Legacy technology requires 40 to 100 seconds to process a document, while D2L accomplishes it in under 1 second.
Pushing Boundaries: It allows models to handle text up to four times longer than their original context window, achieving near-perfect accuracy in "needle-in-a-haystack" retrieval tests.
Text-to-LoRA: Tailor AI with Everyday Language
Text-to-LoRA makes models more responsive. Users simply describe a task in natural language—for example, "help me solve a complex math competition problem"—and the system automatically generates a dedicated plugin to enhance performance. Experiments reveal that adapters created this way can outperform dedicated models trained from scratch on math and logical reasoning tasks.
Powerful Cross-Modal Tech: Enabling Text Models to "See" Images
Researchers discovered an unexpected bonus: D2L exhibits strong cross-modal abilities. By mapping visual information into the parameters of a pure text model, a model that has never processed images before can classify them with an accuracy of **75.03%**.
Sakana AI’s accomplishments significantly lower the barrier for individuals and companies to customize private AI models. They also pave a new way toward developing more lightweight and intelligent general artificial intelligence (AGI).
Paper: https://arxiv.org/pdf/2602.15902
Related article
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
Related Special Topic Recommendations
Comments (0)
0/500
Large language models' struggle with processing long texts, often called "memory anxiety," may soon be a thing of the past. Recently, Tokyo-based AI startup Sakana AI unveiled two breakthrough technologies: Text-to-LoRA (T2L) and Doc-to-LoRA (D2L). Utilizing an innovative "super network" architecture, these technologies enable large models to "absorb" massive documents or learn new tasks in under a second, all without retraining.

AI developers have long faced a difficult trade-off: cramming lengthy documents into a chat window—which slows responses and consumes high memory—or paying the high cost of fine-tuning a model. Sakana AI introduces a third option. Through a one-time payment for pre-training, it generates minuscule weight plugins (LoRA) to achieve low-cost, efficient model adaptation.
Doc-to-LoRA: Memory Needs Drop from 12GB to Just 50MB
This represents the most impressive technology in the latest release. Processing a 128,000-token document (roughly 100,000 words) with traditional methods demands over 12GB of VRAM to store information. With D2L, the model can directly "digest" this information into a plugin smaller than 50MB.
Remarkable Speed: Legacy technology requires 40 to 100 seconds to process a document, while D2L accomplishes it in under 1 second.
Pushing Boundaries: It allows models to handle text up to four times longer than their original context window, achieving near-perfect accuracy in "needle-in-a-haystack" retrieval tests.
Text-to-LoRA: Tailor AI with Everyday Language
Text-to-LoRA makes models more responsive. Users simply describe a task in natural language—for example, "help me solve a complex math competition problem"—and the system automatically generates a dedicated plugin to enhance performance. Experiments reveal that adapters created this way can outperform dedicated models trained from scratch on math and logical reasoning tasks.
Powerful Cross-Modal Tech: Enabling Text Models to "See" Images
Researchers discovered an unexpected bonus: D2L exhibits strong cross-modal abilities. By mapping visual information into the parameters of a pure text model, a model that has never processed images before can classify them with an accuracy of **75.03%**.
Sakana AI’s accomplishments significantly lower the barrier for individuals and companies to customize private AI models. They also pave a new way toward developing more lightweight and intelligent general artificial intelligence (AGI).
Paper: https://arxiv.org/pdf/2602.15902
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ





Home






