注册并分享邀请链接,可获得视频播放与邀请奖励。

Sumanth 的个人资料封面
Sumanth 的头像

Sumanth (@Sumanth_077)

@Sumanth_077
0 正在关注    0 粉丝
I built a self-directed data analyst! 100% Open Source The agent investigates a given dataset on its own without requiring you to guide it through every step. Give it a dataset and an objective like “Why did revenue decline?”, and it decides what to inspect, which analyses to run, what hypotheses to test, and what to investigate next based on the evidence it finds. It can inspect the data, run SQL, use Python for deeper analysis, create charts, form hypotheses, test them against the dataset, and keep following new evidence until it has enough to explain what happened. There is no fixed sequence telling it what analysis to run next. Each result becomes context for the next decision. If a query fails, the agent can use that error to adjust its approach. If it starts repeating similar analyses without making progress, the harness can detect that and push it to reassess the investigation. It also keeps explicit state across the run, including findings, hypotheses, tool history, and usage, instead of relying only on raw conversation history. For the model layer, I used Liner Mark 1.0. This works well for this kind of setup because one investigation can involve many model calls, but every step does not need the same level of reasoning. Liner Mark 1.0 routes each request to an appropriate underlying model while exposing a single model interface to the agent. So the same loop can move between dataset inspection, query planning, hypothesis evaluation, error recovery, and final synthesis without manually choosing a different model for every step. The workflow looks like this: Dataset + Objective → Investigate → Run SQL/Python → Observe → Update Hypotheses → Decide Next Step → Repeat → Final Report At the end, the agent returns the root cause, supporting findings, hypotheses it tested, relevant charts, recommended next steps, and a usage breakdown. GitHub repo: What I like about this setup is that the analysis path is not predetermined. You start with a goal, inspect the evidence, form a theory, test it, and let the results decide what to investigate next.
显示更多
Hands on AI Engineering! I open-sourced a collection of 50+ hands-on AI engineering tutorials. It features step-by-step projects and tutorials on: • AI Agents and Multi-agents • RAG (Agentic, Vision, and Local) • MCP AI Agents • OCR Apps • Voice AI Agents • & so much more 100% free and open source. 1k+ Github stars I've shared the link in the comments!
显示更多
Stop generating text when all you need is a decision! Jev is TypeSafe AI’s first “System One” model, built for a part of AI systems that we currently keep forcing generative LLMs to handle: making bounded decisions inside software. Think about what happens during an agent run. The main model may be writing code, researching, planning, or reasoning through a problem, but the system around it is constantly making smaller decisions. Which model should handle the next step? Does this tool call need approval? Is the agent still making progress? Is the task actually complete? Today, we often send those questions back to another generative LLM. Even with structured outputs, the model is still generating tokens sequentially and constraining them into the schema we asked for. The application then takes that generated output and turns it back into the decision it needed in the first place. Jev removes string generation from that loop. You give it the current state and define the decisions you care about. Those can be a yes/no judgment, a choice between predefined options, or a score across an ordered scale. Jev returns typed answers with probabilities that the application can use directly. That distinction becomes useful when these decisions sit everywhere inside an agent harness. A coding agent can use the main LLM to understand a repository and write a fix, while Jev handles model routing, tool gating, risk checks, progress evaluation, or deciding whether another iteration is actually necessary. And Jev should not become the permission system either. A probability that an action is safe is still a model judgment. Permissions, spend limits, allowlists, and other hard constraints should remain deterministic. The architecture I find more interesting is the separation itself: generative models do the open-ended work, decision models handle fuzzy but bounded judgments around that work, and the runtime decides what is actually allowed to happen. As agents become longer-running, the number of these small decisions only increases. Using a full generative model for every one of them may turn out to be a very expensive abstraction. If you want a deeper look at Jev, I also wrote a separate breakdown of how it works. I’ve quoted the article below.
显示更多
I built a self-evolving code review agent! Most code review agents use the same prompt every time they run. They can review a diff and generate comments, but they do not remember what your team accepted, rejected, or corrected in previous reviews. That means the same mistakes can keep showing up across review cycles. This project adds persistent experiential memory to the review loop. Before every review, the agent retrieves relevant team rules and similar past review trajectories from memory, then uses that context alongside the current diff to generate review comments. After the review, the engineer can accept, reject, or edit every comment. That feedback becomes the learning signal. Accepted comments reinforce useful rules. Rejected comments create lessons about what should not be flagged. Edited comments refine the team’s preferred convention, scope, or wording. The agent adapts non-parametrically. Instead of fine-tuning the underlying model, it improves by storing review outcomes as reusable memory and retrieving the most relevant experience before the next review. Here is how the loop works: • Retrieve: search memory for relevant insights and similar past reviews before generating comments • Review: generate structured feedback using the current diff plus retrieved memory • Human feedback: pause the workflow and collect an accept, reject, or edit decision for every comment • Reflect: convert those decisions into reusable natural-language rules with rationale, polarity, scope, and confidence • Persist: store both the learned insights and complete review trajectory for future retrieval Over time, the agent builds a memory of how your team actually reviews code without retraining the underlying model. The full workflow runs locally with LangGraph, Ollama, BGE embeddings, and a vector database for long-term memory. The interesting part is that every review leaves behind experience that can be retrieved and reused on the next one. Github Repo:
显示更多
Hands on AI Engineering! I open-sourced a collection of 50+ hands-on AI engineering tutorials. It features step-by-step projects and tutorials on: • AI Agents and Multi-agents • RAG (Agentic, Vision, and Local) • MCP AI Agents • OCR Apps • Voice AI Agents • & so much more 100% free and open source. 1k+ Github stars I've shared the link in the comments!
显示更多
ByteDance dropped a banger paper on self-evolving agent harnesses! HarnessDev evaluates whether AI agents can build a runnable harness from scratch and iteratively improve it using execution feedback. Most agent benchmarks keep the harness fixed and only evaluate the model inside it. HarnessDev changes the target of evaluation itself. The agent starts from a minimal seed, builds the harness around the task, runs it, observes what worked or failed, and then modifies that harness across multiple iterations. That means the agent is not only solving the task. It is also changing the planning, memory, tool use, state management, and execution logic around itself. The paper evaluates this in two stages: • Creation: can the model build a complete runnable harness from a minimal starting point? • Evolution: can it improve that harness using feedback from previous runs? The interesting part is that runnable does not automatically mean better. Some generated memory and state mechanisms existed in the code but were barely used during execution, and improvements on visible feedback did not always transfer to held-out tasks. Only 34 of 64 harness changes moved in the same direction on both visible feedback and held-out evaluation, and only 2 of 9 final harness versions were actually the best-performing version on the held-out set. So the paper is really exposing a new challenge: Agents can already start modifying the infrastructure they run on. The harder part is making sure those changes actually generalize. I've shared the paper in the comments!
显示更多
Loop vs Graph Engineering: Clearly Explained! A loop is an autonomous cycle where an agent plans, acts, verifies the result, and keeps iterating until it reaches a stopping condition. The basic pattern is simple: Plan → Act → Verify → Repeat Loop engineering is about making that cycle reliable. You decide what context survives between iterations, how the agent evaluates its own work, when it should retry, how failures are handled, and what condition finally ends the run. This works well when one agent can own a coherent task from start to finish. The problem starts when the work itself becomes more structured. Some tasks need to happen in parallel. Different agents may need different context, tools, memory, or permissions. Certain steps may require human approval, while others can continue automatically. At that point, trying to force everything through one loop makes the system harder to reason about. That is where graph engineering becomes useful. A graph makes the workflow explicit. Nodes represent units of work such as agents, tools, validators, or deterministic functions, while edges define how state and control move between them. The important part is that loops and graphs are not competing architectures. A loop can itself be one node inside a larger graph. For example, a research agent may run its own Plan → Act → Verify loop, while a separate analysis agent runs another loop. A graph can then coordinate when those agents start, what context they receive, how their outputs are combined, and what happens next. So the practical distinction is: • Loop Engineering: design one autonomous execution cycle around a task • Graph Engineering: design how multiple execution cycles, tools, and decision points coordinate through shared state and explicit control flow You usually start with a loop when one agent can handle the task end to end. You move to a graph when the work needs parallel execution, specialized agents, different permissions, explicit branching, or clearer control over how state moves through the system. Graph engineering does not replace loop engineering. It becomes useful when one loop is no longer enough to hold the workflow together. I've shared the full article in the comments!
显示更多
Your AI agent doesn’t need a better model. It needs a better harness! JIT-Agent is a compact meta-agent that writes your agent harness on the fly. Most agent systems use a fixed scaffold for every task. The same planner, the same memory setup, the same tool orchestration, and the same execution strategy regardless of what the agent is actually trying to solve. JIT-Agent changes that by generating a task-specific harness across four core modules: memory, planning, action, and capability orchestration. So instead of treating the harness as static infrastructure, it becomes something the agent can compose based on the task itself. It can also improve that harness from execution traces and feedback without updating the generator model. This matters because a large part of agent performance comes from everything around the model. How context is stored, how the task is decomposed, which tools are available, when they are called, and how the execution loop is structured can change the final result significantly. The benchmarks make that pretty clear. With JIT-Agent, DeepSeek-V4-Flash outperformed GPT-5.6 by 9.1 points on DeepSearchQA and 4.3 points on OdysseyBench. GLM-5.2 also improved by up to 20.2 points with a better generated harness. Same model family. Better harness. Better agent. That is the part I find most interesting. As agent systems get more complex, improving the model may not always be the highest-leverage move. Improving the harness around it might matter just as much. 100% open source. I've shared the GitHub repo and paper in the replies!
显示更多
The real bottleneck in agentic coding isn’t the model. It’s how long you can keep the loop running. Modern coding agents work through loops. They plan, write code, inspect the result, fix errors, test again, and keep iterating until the task is complete. Every pass through that loop means more model calls and more usage. That creates a simple problem. The more you rely on the agent to think, explore, and iterate, the more expensive the workflow becomes. Replit’s new Free Mode is designed around this. Replit Agent already lets you start with an idea, build something, ask follow-up questions, change direction, and keep iterating. What changes now is how much of that loop you can run without constantly thinking about usage. Instead of trying to squeeze everything into one perfect prompt, you can keep the agent in the loop for longer. Ask a question. Explore an approach. Build it. Inspect the result. Change direction. Keep going. The new conversation experience also keeps that entire process in the same context, so thinking through the idea and actually building it no longer feel like separate workflows. This is the part I find interesting. As coding agents become more iterative, the constraint is no longer just model quality. It is also how much room you have to keep the loop running. Replit is reducing that friction with Free Mode. I've shared the link in the replies!
显示更多
Replit Free Mode, powered by @OpenAI GPT-5.6 Luna. Let’s make intelligence accessible to everyone.
NVIDIA open-sourced the Pythonic way to build AI agents! NOOA (NVIDIA Object Oriented Agents) is a Python framework that collapses the separate abstractions most agent frameworks use into a single Python class. No separate prompt templates, tool schemas, or callback wiring. The agent is just a Python object. The design maps directly to Python concepts you already know. State lives as typed class fields. Capabilities are methods. Docstrings are prompts. Type annotations are contracts. Methods with a real body stay deterministic Python. Methods with `...` as the body become LLM-driven at runtime - the model implements them on the fly. The model acts by writing Python in a REPL with access to `self`, imports, and helpers. Python methods and type annotations supply the callable interfaces, so there's no need to write separate tool schema definitions. The method name, parameters, and docstring are the prompt. Because agents are just Python classes, you can test them with pytest, trace them, refactor them, and version control them the same way you treat the rest of your software. One honest note from the README: NOOA is research software. LLM-generated code can take dangerous actions. Run agents that execute generated code in a sandboxed environment, not directly on your filesystem. Key capabilities: • Agents as Python classes: state, methods, prompts, and type contracts in one place • Methods with real bodies stay deterministic, methods with "..." become LLM-driven • LLM acts via Python REPL with access to self and imports • Built-in tracing for every LLM call, code execution, and method invocation • Long-term memory subsystem via nooa-memory • MCP support and Harbor benchmark runner I've shared the link in the replies!
显示更多
0
12
135
29
转发到社区
Cloudflare open-sourced their own internal AI OS! Cloudflare OS is an AI productivity workspace originally built for internal use at Cloudflare. A large portion of their workforce - engineering, sales, and everything in between - uses it every day. They're open-sourcing it so others can fork and customize it as their own company OS. The core idea is a departure from how cloud software has worked for the past 25 years. When you create a slide deck in Cloudflare OS, you're not connecting to a shared SaaS app running on someone else's server. The system creates a private instance of slide deck software just for you, running in its own sandbox. They call these Gadgets. This has two direct consequences. First, a security bug in the slide deck app can't leak your slides to anyone else because your instance is completely isolated. Second, if the app is missing a feature you need, you can ask the agent to add it. And because you're running your own copy, it's safe to do so. The security model underneath this is called Gatekeepers. Each external resource connection gets a Gatekeeper that handles authorization, enforces narrow access to only what you intended, and logs every action for review. The genuinely interesting part is how Gatekeepers handle human approval. Most agent setups stop and wait for the human to approve each action before continuing - which is why people often end up using "dangerously-skip-permissions." Gatekeepers simulate the action locally instead, let the agent keep working, queue the real action, and let the human approve or reject in bulk later when convenient. Key capabilities: • Gadgets: private per-user instances of every app, sandboxed and AI-modifiable • Gatekeepers: async human-in-the-loop approval without blocking agent progress • Blueprints: shareable app templates where each user runs their own copy • Built-in coding agent that builds, tests, and debugs Gadgets • Real-time multiplayer collaboration via Durable Objects • Capability-based security - agents get access to nothing by default • Runs on Cloudflare Workers or self-hosted on workerd I've shared the link in the replies!
显示更多
The distributed platform that powered Kimi K3's RL training! AgentENV is the infrastructure that powered agentic RL training for Kimi K3 - running thousands of isolated sandboxes simultaneously, each forkable, snapshotable, and resumable in milliseconds. Training agents with RL means running the same task across thousands of parallel environments. Each needs its own isolated sandbox where the agent can write code, run shell commands, and interact with the filesystem. Docker is too slow to start. Full VMs are too heavy. And at training scale, cloud sandbox costs compound fast. AgentENV uses Firecracker microVMs. Environments boot or resume in under 50ms and pause in under 100ms. When an agent finishes its turn and waits for the next update, the environment pauses and returns its memory to the host. When work arrives again, it resumes instantly from exactly where it left off. The fork capability is what makes parallel RL training practical. Instead of booting thousands of fresh VMs from the same base state, AgentENV snapshots one running environment and forks it into multiple independent sandboxes in under 100ms. Each fork is fully isolated. Agents try different approaches simultaneously without interfering with each other. Snapshots happen incrementally, completing in under 100ms even under heavy disk modification. They persist to S3-compatible object storage so no state is lost if a machine goes down. Local disk acts as a bounded cache, so images can exceed disk capacity without pre-warming every host. AgentENV also exposes an E2B-compatible HTTP API. If your agent code already uses the E2B SDK, point one environment variable at your AgentENV server and your existing code works without any changes. Key capabilities: • Firecracker microVM environments: boot and resume in under 50ms • Fork a running environment into multiple independent sandboxes in under 100ms • Incremental snapshots to S3-compatible storage in under 100ms • Memory ballooning returns idle guest memory to the host • Images can exceed disk capacity via overlaybd with on-demand loading • E2B-compatible HTTP API - drop-in replacement with no code changes • Distributed across machines via Kubernetes or Docker Compose I've shared the link in the replies!
显示更多
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: Tech report: Tech blog:
显示更多
Andrew Ng open-sourced an AI coworker! OpenWorker is a desktop app that delivers finished work, not just chat. You describe the outcome you want - a customer brief, a triaged inbox, a calendar update, a release status across Jira and GitHub - and it works across your files, terminal, and connected apps to produce the actual deliverable. The design principle: it reads freely but writes only after you approve. Before anything consequential - sending a message, making a calendar change, running a shell command - it pauses and shows you exactly what it's about to do. Unattended scheduled runs park approval-needed actions in an inbox rather than acting on their own. 25+ integrations out of the box: GitHub, Slack, Jira, Notion, Linear, HubSpot, Gmail, Google Calendar, Outlook, monday, and your terminal and local files. Any MCP tool plugs in too. Works from Slack directly. Mention @ OpenWorker in a channel, the work happens on your desktop with your tools, and the answer comes back as a thread reply. Bring your own model. OpenAI, Anthropic, Google Gemini, DeepSeek, Kimi, Qwen, MiniMax, Mistral, Grok, and fully local via Ollama. Everything runs on your machine with your own API keys. Built on aisuite, Andrew Ng's own unified LLM provider library. Currently in open beta on macOS and Windows. Key capabilities: • Delivers finished deliverables: documents, reports, Slack replies, updated calendars • Approval-gated writes, sends, and shell commands • 25+ integrations including GitHub, Slack, Jira, Notion, Gmail, and MCP tools • Scheduled automations for recurring work • Works from Slack via @ OpenWorker mention • Bring your own model: any provider or fully local via Ollama • Local-first: everything runs on your machine OpenWorker shows what happens when you get the loop right. Define the goal, the agent executes autonomously, approval gates control what runs. That's loop engineering in practice. Wrote a full breakdown on how to design loops you can actually trust to run without you. I've shared the link to the Github Repo in the replies!
显示更多
Stop optimizing tokens. Index your context instead! Retrieval quality is the foundation of context engineering. Karpathy described it best: the heavy cognitive work should happen at ingestion, not at query time. When knowledge is properly structured before retrieval, the model's job becomes reasoning, not sorting. Most AI systems focus on compressing what the model sees. The more important problem is what gets retrieved before the model sees anything. Token efficiency starts at retrieval. When context is properly indexed and prepared, the model spends its tokens on reasoning. When retrieval is weak, the model spends those same tokens sorting through noise and filling gaps from its own weights. That's the silent failure mode. Retrieval returns topically correct but incomplete context. The model completes the gaps from parametric knowledge and streams it out the same way as grounded content. No signal in the output tells you which parts came from retrieved context and which came from weights. The answer looks confident. It just isn't complete. The common assumption is that hallucination is the main failure mode. It's not. Models handle off-domain questions reasonably well now. If nothing in the retrieved context looks relevant, there's no material to build an answer on. The harder failure is partial coverage. The right document was retrieved. But not the full picture. Coverage gaps don't produce error messages. They produce confident answers with pieces missing. This gets worse when sources stay isolated. The same person might appear across multiple tools and systems. If those sources are indexed separately, the model has to figure out they refer to the same entity on its own. That's work that should happen before the model starts reasoning. Glean's system of context is built around this problem: • Unified index across all connected applications, not each source kept separate • Specialized indexes for different types of information: company data, code, experts, profiles, tools, and calendars • Multiple retrieval methods - semantic when meaning matters, lexical when exact terms matter, structured when fields and relationships need to stay intact • Enterprise Graph that maps relationships across people, teams, customers, and projects so relevance reflects how the company actually works • Memory that carries context forward across sessions and tasks • Tools that let the model act on what it finds The gap between finding information and understanding it is where most AI systems fall short. I've shared the link in the replies!
显示更多
YC just open sourced their multi-agent harness! QM is the multiplayer agent harness Y Combinator built and has been running internally across accounting, legal, events, and engineering. They used QM to build QM itself. The starting point was a fleet of 50+ Hermes agents - one personal assistant per employee. Managing that many separate agents became complex. QM came from asking a different question: instead of one agent per person, what if a company had one harness that worked for everyone? Most agents are designed like personal assistants. QM is designed for teams. Each employee gets their own isolated workspace with scoped memory, files, permissions, crons, and a durable sandbox. Those workspaces also connect in shared Slack channels and projects where people and the agent work together. The same identity and configuration carries between Slack and the web app. Admin controls which harnesses and models are available org-wide. Pi, OpenCode, Codex, and Claude Code all drive the same core, so a deployment isn't tied to any single vendor. Background work runs while nobody's watching. Crons and webhooks trigger tasks across the org. Skills are scope-owned and shareable by grant, with admin-gated promotion to the whole organization. YC is direct about where it stands: it's an experiment, it's early, and it has bugs. Key capabilities: • Multiplayer: personal workspaces + shared Slack channels and projects • Works natively in Slack and on the web • Vendor agnostic: Pi, OpenCode, Codex, Claude Code all supported • Per-scope memory, files, keychain, permissions, crons, and durable sandbox • Background crons and webhook triggers • Shareable skills with admin-gated org promotion • Three security postures: Strict, Auto, Dangerous 100% open source. The multiplayer harness problem is real and QM is a solid approach to it. The other harness problem most teams haven't named yet is context quality - what actually flows into the harness from your organization's systems determines everything that follows. Wrote a detailed breakdown on exactly that. I've also shared the link to QM in the replies!
显示更多
RAG system that skips HTML parsing entirely! PixelRAG is an open-source visual RAG framework that renders documents as screenshots instead of parsing them into text. Most RAG pipelines start by converting HTML to text. Tables flatten into unstructured rows. Charts disappear. Layout context is gone before the LLM ever sees it. The paper measured this directly: HTML-to-text conversion accounts for 36.6% of retrieval failures on SimpleQA. PixelRAG skips that step entirely. It renders pages as screenshot tiles using Playwright, embeds those tiles with a fine-tuned Qwen3-VL-Embedding model, builds a FAISS index, and passes retrieved images directly to a VLM reader. No text abstraction in between. Benchmarked across six datasets against the strongest text-based baselines: - SimpleQA: 78.8% vs 71.6% (+7.1 points) - NQ-Tables: 48.8% vs 42.5% (+6.3 points) - EVQA: +15.5 points - LiveVQA: +11.3 points One honest caveat from the authors: this requires Qwen3-VL-4B class models or larger to see the benefit. Smaller models trail text retrieval. The authors also recommend using PixelRAG as an enhancement layer alongside existing text systems rather than a full replacement. Ships with a pre-built Wikipedia index covering 8.28M articles across 28.1M screenshot tiles. A Claude Code plugin lets Claude take screenshots of any URL and reason over the visual content directly. Key capabilities: • Renders web pages, PDFs, and images as screenshot tiles via Playwright • Fine-tuned Qwen3-VL-Embedding model for visual retrieval • FAISS index for fast vector search • Pre-built Wikipedia index: 8.28M articles, 28.1M tiles • 3x token cost reduction via image compression • Claude Code plugin for direct URL screenshot and visual reasoning • LoRA fine-tuning support via pixelrag-train 100% open source. I've shared the link in the replies!
显示更多
Turn any website into agent-ready data! Loop engineering is about designing systems that run agents autonomously. Instead of prompting your agent manually each turn, you write a loop that finds the work, hands it to the agent, checks what came back, and decides what happens next. Your job is to design the loop once and walk away. But loops that need live information from the web hit a real constraint. JS-heavy pages return empty content. Anti-bot systems return challenge pages. Login walls block access entirely. When the model gets weak context back, it still responds - just less accurately. Anakin is building the source-access layer underneath agents. URL Scraper turns any URL into clean Markdown, HTML, or structured content immediately usable by an LLM. Built for scale across 200M+ active websites globally, including a large chunk of Cloudflare and Akamai-protected pages. Authenticated sessions handle content behind login walls. Wire handles workflow-heavy sites. Login flows, navigation, form submission, report exports - all accessible through a stable API. Define the workflow once and Wire keeps it working as websites change. Key capabilities: • Clean Markdown, HTML, or structured output from 200M+ active websites • Built for difficult pages including Cloudflare and Akamai-protected sources • Authenticated sessions for content behind login walls • Wire for login flows, navigation, form submission, and export-based access • Useful for AI agents, RAG systems, finance intelligence, and vertical AI workflows I've shared the link in the replies!
显示更多
Stop prompting AI agents. Design the loops that prompt them instead. This is the core idea behind Loop Engineering, a methodology for building automated systems that orchestrate your AI coding agents instead of prompting them manually. Boris Cherny, Head of Claude Code at Anthropic, puts it directly: "I don't prompt Claude anymore. I have loops running that prompt Claude and figure out what to do. My job is to write loops." A loop is an automated pipeline that runs on a schedule. It checks what needs to be done, prompts your AI coding agent with the right context, verifies the result, and either commits the fix or escalates to you. Then it runs again. This repo is a starter kit and reference guide for building these loops. Seven production-ready patterns, CLI tools to scaffold your setup, and documentation covering failure modes, anti-patterns, safety, and multi-loop coordination. The seven patterns: Daily Triage, PR Babysitter, CI Sweeper, Dependency Sweeper, Changelog Drafter, Post-Merge Cleanup, and Issue Triage. Each one ships with a starter kit, cadence recommendation, and token cost estimate. Three CLI tools handle setup. "loop-init" scaffolds skills, state, and budget files and prints your Loop Ready score. "loop-audit" checks how ready your setup is and suggests improvements. "loop-cost" estimates token spend before you run anything. Works with Claude Code, Codex, Grok, OpenCode, Cursor, and GitHub Actions. Key capabilities: • 7 production loop patterns with starters and token cost estimates • loop-init scaffolds your setup and prints a Loop Ready score • loop-audit scores readiness and suggests improvements • loop-cost estimates token spend per cadence • Failure modes, anti-patterns, and safety documentation included • Works with Claude Code, Codex, Grok, OpenCode, Cursor 100% open source. I've shared the link in the replies!
显示更多