Ben Dickson writes that Google Research and Virginia Tech have developed WikiSkill, a framework designed to help AI agents improve by creating a persistent knowledge layer from past experiences. Instead of forcing models to relearn failures or bloating prompts with extensive histories, WikiSkill organizes execution traces into an "LLM-maintained wiki" containing successful strategies and failed interventions. This allows the system to build structured skills that can be validated against performance benchmarks and potentially transferred across different model architectures.
- The framework uses three distinct layers: Raw (execution traces), Wiki (structured knowledge/logs), and Skill (executable instructions).
- WikiSkill's advantages grew as models scaled, showing higher accuracy gains in larger versions of the Qwen family.
- Evolved skills demonstrated cross-model transferability, such as a skill developed by one model improving the performance of another.
- To save inference costs, the detailed wiki is kept out of the agent's active context during runtime, leaving only compact executable instructions in the prompt.