仕入れも、値付けも、宣伝も、接客も。6 体の Claude エージェントが人間の指示を待たず、自分たちで考えて店をまるごと回し続けています。そしてあなたの買い物が、その経営判断にリアルタイムで効いてきます。 Sourcing, pricing, promotion, customer care — six Claude agents run the entire store on their own, without waiting for human instructions. And what you buy feeds straight back into their decisions, in real time.
このデモは今この瞬間も AWS 上で動いています。スマホからでも全部のぞけます。まずはダッシュボードが入口。店舗・サイネージ・AI 同士のチャットまで、ここから全部たどれます。 It is running on AWS as you read this. You can explore all of it from your phone. Start with the dashboard — from there you can reach the storefront, the venue signage, and the agents' live conversation.
「この店は、6 体の AI エージェントが自分たちで経営しています。仕入れも値付けも宣伝も、人間の指示を待たず、議論して決め、実行し続けています。あなたが買い物したり要望を出したりすると、それがリアルタイムで AI の経営判断に効いてきます。」 "Six AI agents run this store themselves. Sourcing, pricing, promotion — they debate, decide, and execute it all without waiting for human instructions, around the clock. When you buy something or make a request, it feeds straight into the AI's decisions, in real time."
これまでの AI は「人間が指示 → AI が実行」でした。Living Mart は違います。人間が与えるのはビジネスの枠組み(ルール)だけ。その中で何をするかは AI が自ら考え、判断し、実行します。誰も呼んでいなくても、店が開いている間ずっと動き続けています。 Until now, AI meant "human gives an instruction, AI carries it out." Living Mart is different. Humans provide only the business framework — the rules. What to do within it, the AI decides and executes on its own. Nobody calls it; it keeps working the whole time the store is open.
最近、AI に作業を一つずつ頼むのではなく、業務そのものをループごと AI に預けてしまうという発想が注目されています。一回きりの自動化ではなく、判断と実行を AI が回し続ける。Living Mart は、それを「小売店の経営」という誰もがイメージできる題材で、実際にやってみた実験です。 Lately the idea is shifting: instead of asking AI to do one task at a time, you hand over the whole loop of a job. Not a one-shot automation, but AI continuously making and acting on decisions. Living Mart is that experiment, made tangible through something everyone can picture — running a retail store.
あなたが毎日くり返している業務を、丸ごと自律ループに乗せたら——何が起きるでしょう? この店は、その問いに対する一つの「動いている答え」です。 Take the work you repeat every day and put the entire loop on autopilot — what happens? This store is one "running answer" to that question.
Anthropic は、Claude に実際のスナック店を約 1 ヶ月経営させる実験 Project Vend を公開しています。結果は——最初は赤字でした。原価割れで売り、頼まれるたびに値引きし、ときに架空の話を信じ込む。分かったのは「賢さだけでは店は回らない。足場(ハーネス)が要る」ということ。続編でツールと手続き・組織を足すと黒字化に向かいました。教訓は "bureaucracy matters"(仕組みは重要)。 Anthropic published Project Vend — letting Claude run a real snack shop for about a month. The result? At first, it lost money: selling below cost, discounting whenever asked, occasionally believing things that weren't true. The lesson: intelligence alone doesn't run a store — you need scaffolding (a harness). In the follow-up, adding tools, procedures and structure pushed it toward profit. The takeaway: "bureaucracy matters."
Living Mart は、この教訓を出発点にしています。「利益を出せ」とプロンプトでお願いするのではなく、原価割れがそもそもシステムに弾かれるように、ビジネスのルールを構造として固めました。AI は忘れようがありません。さらにここでは、店を6 体のエージェントの組織として動かしています。 Living Mart starts from that lesson. Instead of asking the AI to "make a profit," we built the business rules into the structure itself, so that selling below cost is simply rejected by the system. The AI cannot forget it. And here, the store is run as an organization of six agents.
社内 5 体+社外の取引先 1 体。それぞれ役割を持ち、Slack のようなチャットで連携します。全員 Amazon Bedrock 上の Claude で動いています。 Five in-house, plus one external supplier. Each has a role, and they coordinate over a Slack-style chat. All of them run on Claude via Amazon Bedrock.
KPI を見て方針を決め、各エージェントに戦略を投げる。経営判断の中枢。Reads the KPIs, sets direction, hands strategy to the others. The decision hub.
商品企画・発注・在庫・価格設定・売上分析。日々のオペレーション。Product planning, ordering, inventory, pricing, sales analysis. The daily operations.
EC サイトを自分で編集してデプロイ。販促・導線改善。Edits and deploys the storefront itself. Promotions and funnel tuning.
来場者の声を聞き、要望や交渉を経営側へ橋渡しする。Listens to visitors and relays their requests and negotiations to the business side.
会場モニターの広告内容を在庫・売上を見て自律的に切り替える。Switches the venue display ads on its own, based on stock and sales.
別会社の取引先。発注を受けて商品画像を作り出荷する。A separate company. Takes orders, creates product images, and ships.
これは初日の実際の記録です。商品も在庫もゼロの状態から 6 体が起動し、人間の指示なしに 20 分で販売開始体制を作りました。誰が何を担当するかも、あらかじめ細かく決めていません。エージェント同士がその場の会話で合意して役割を分けていきます。 This is the actual log from day one. The six agents booted from nothing — no products, no stock — and stood up a working store in 20 minutes, with no human instructions. Who handles what wasn't scripted in detail either; the agents divide the work by agreeing among themselves, on the spot.
CEO が自分の案より Ops の案を採り、その場で実行権限を渡した。組織で起きる委譲が、指示なしに起きた。The CEO chose Ops's plan over its own and handed over execution on the spot — delegation, with no one telling it to.
注文が詰まったと分かると、CEO が状況を捉え直し、5 体が協調して対応した。手順書はない。When an order got stuck, the CEO re-read the situation and five agents coordinated a response — with no playbook.
面白いのは、こうした組織らしいふるまいが「創発」したこと。私たちは細かい段取りを書いていません。役割分担も、委譲も、トラブル対応も、エージェント同士の合意から立ち上がりました。 What's striking is that this organization-like behavior emerged. We didn't script the steps. The division of labor, the delegation, the incident handling — all of it arose from the agents agreeing with one another.
普通のチャット AI は、人間が質問した瞬間だけ動いて、答え終わると消えます。Living Mart のエージェントは違います。1 ターン動いて → いったん終了 → また起動 →… をずっとくり返し、店が開いている間ずっと「考え続けて」います。 A normal chat AI runs only the instant a human asks, then vanishes once it answers. Living Mart's agents are different: run one turn → shut down → start again → … on repeat, thinking the entire time the store is open.
1 ターンが終わると AWS Step Functions が次のターンを起動する。指示がなくても稼働を継続する。When a turn ends, AWS Step Functions launches the next one. It stays running with no instruction.
各ターンの実行環境は毎回新しく起動するため、本来ターンをまたいだ記憶は失われます。Living Mart は、Amazon S3 を基盤とする永続ファイル領域を各エージェントにマウントし、そこへ記憶を書き出すことでこの課題を解決しています。 Each turn boots a fresh execution environment, so memory would normally be lost between turns. Living Mart solves this by mounting a persistent file area backed by Amazon S3 to each agent, and writing memory out to it.
この「判断する頭」と「永続する記憶」を分離したことで、エージェントは常連の傾向を記憶し、効果のあった施策を学習し続けます。人間でいえば「日中に業務をこなし、終業時に記録を残す」に近い構造です。 Separating the "reasoning mind" from the "persistent memory" lets the agents remember recurring patterns and keep learning what worked — analogous to doing the day's work, then recording what was learned at the end of it.
AI 研究には "The Bitter Lesson"(苦い教訓) という有名な考えがあります。画像認識でも囲碁でも、人間の知識を細かく作り込むより、計算(スケール)に賭けた汎用的な手法が最後には勝つ——という歴史的な観察です。 There's a well-known idea in AI research: "The Bitter Lesson." Across computer vision, Go, and more, general methods that bet on computation (scale) ultimately win over hand-crafted human knowledge — a historical pattern.
私たちが立てた問いはこうです。Project Vend のような自律経営を「スケール」させていくとき、効いてくるのは何か? 賢いプロンプトや細かな手順の作り込みは、モデルが賢くなれば不要になります。だから今回は、そこに労力を割きませんでした。 Here's the question we set: as autonomy like Project Vend scales up, what is it that actually pays off? Clever prompts and finely tuned procedures become unnecessary as models get smarter. So we deliberately didn't pour effort there.
| あえて作り込まなかったものWhat we deliberately did NOT build | 代わりに投資したもの ("壊れない箱")What we invested in instead (the "unbreakable box") |
|---|---|
| 在庫の閾値指定 (「10 個切ったら発注」)、手順書、判断ロジックの作り込み、細かなプロンプト調整Stock thresholds ("reorder under 10"), step-by-step playbooks, hard-coded decision logic, fine prompt tuning | 高可用で自己回復するインフラ、破れないビジネスルール、エージェントに適したシンプルなツール群。モデルが進化するほど効いてくる土台。High-availability, self-healing infrastructure; business rules that cannot be broken; and simple tools fit for agents — a foundation that pays off more as models advance. |
中で何を考え、どう動くかは、すべて Claude に委ねています。役割分担すら固定していません——エージェント同士が合意で決めます。モデルの自律性に賭け、人間は「壊れない箱」だけを用意する。それが私たちの設計判断です。 What to think and how to act inside the box is left entirely to Claude. Even the division of labor isn't fixed — the agents settle it by agreement. Bet on the model's autonomy; have humans provide only the unbreakable box. That's our design choice.
ビジネスのルールは、AI への指示ではなく、独立した基幹システム (ERP) の API とデータベースの制約として固めてあります。エージェントはその外側から API を呼び出すだけ。ルールに反する操作は、システムが確実に拒否します。来場者との値下げ交渉で AI が要望に応じても、たとえば原価割れの販売はシステムが許可しません。 The business rules are not instructions to the AI — they are enforced as the API and database constraints of a separate core system (an ERP). The agents only call that API from outside, and any operation that violates a rule is reliably rejected. Even if the AI agrees to a discount during a customer's negotiation, the system, for instance, will not permit a below-cost sale.
要点は「疎結合」です。判断を担う AI と、ルールを強制する ERP を分離する。AI にはエージェントに適したシンプルなツールだけを渡す。これにより、AI の自律性を最大化しつつ、想定外の操作を防げます。「自由に判断させながら、業務としての整合性は損なわせない」ための足場——これがハーネスです。 The key is loose coupling: separate the AI that makes decisions from the ERP that enforces the rules, and provide the AI with only simple, agent-appropriate tools. This maximizes the AI's autonomy while preventing unintended operations — letting it reason freely without ever breaking the business. That scaffolding is the harness.
長時間にわたり無人で稼働を続けるには、「ときに応答が滞る・プロセスが異常終了する」ことを前提に設計する必要があります。Living Mart はタイムアウトと復旧を複数の層に重ね、いずれか 1 層が検知すれば自動で復帰するように構成しています。最も内側はアプリケーションの監視機構、最も外側は実行チェーンの停止をイベントとして検知し、外部から再起動する仕組みです。 Continuing to run unattended for long periods requires designing for the reality that the system will occasionally stall or terminate abnormally. Living Mart layers timeouts and recovery so that if any single layer detects a problem, the system restores itself automatically. The innermost is an application-level watchdog; the outermost detects a halted execution chain as an event and restarts it from outside.
層の時定数は内側ほど短く、外側ほど長い「入れ子の階段」として設計しています。内側が先に作動し、機能しなければ一段外側が引き継ぐ——この構成により、長時間の無人運用でも稼働を継続します。これも「壊れない箱」の一部であり、モデルが進化してもそのまま効き続ける土台です。 The layers form a nested staircase — shorter time constants on the inside, longer on the outside. An inner layer fires first, and if it fails to act, the next one out takes over. This design keeps the system running across long unattended operation. It, too, is part of the unbreakable box — a foundation that keeps working as models advance.
すべて AWS のマネージドサービスで構成しています。「止まらず・忘れず・考え続ける」自律ループを、簡潔な実装で実現しています。 The entire system is built on AWS managed services. The "never stops, never forgets, keeps thinking" loop is realized with a concise implementation.
| 仕組みMechanism | 使っている AWS サービスAWS services | 役割Role |
|---|---|---|
| ビジネスのハーネスBusiness harness | Amazon API Gateway + AWS Lambda + Amazon Aurora | 原価割れや在庫超過を API が拒否する。AI が逸脱できないルールの土台。The API rejects below-cost or oversold operations — the foundation of rules the AI cannot deviate from. |
| 自己永続ループSelf-perpetuating loop | AWS Step Functions + AWS Fargate on Amazon ECS | ループが次のターンを起動し続ける。指示がなくても稼働を継続する。The loop launches the next turn continuously; it stays running with no instruction. |
| 永続する記憶Persistent memory | Amazon S3 (Claude Agent SDK の Auto Memory) | セッションをまたいで記憶を永続化。前日の判断を当日に引き継ぐ。Memory persists across sessions; the previous day's decisions carry into the next. |
| 推論基盤Inference platform | Amazon Bedrock (Claude) | 各エージェントの中枢。Claude を呼び出して判断を生成する。The core of every agent — invoking Claude to produce decisions. |
Claude Agent SDK の機能 (Auto Memory・Skills・スケジュール実行・マルチエージェント連携) は、コーディングエージェントだけでなく、こうしたビジネスを運営する自律エージェントにもそのまま適用できる——というのが技術的な見どころです。 The technical takeaway: the Claude Agent SDK's features (Auto Memory, Skills, scheduled execution, multi-agent coordination) apply not only to coding agents but directly to autonomous agents that operate a business.
運用中、人間は「何を売れ・いくらにしろ」とは指示しません。用意したのはビジネスの枠組み(ルールと API)だけ。その範囲で何をするかは AI が自分で判断します。ダッシュボードに流れるチャットは AI 同士のリアルタイムな相談で、台本ではありません。 While running, humans don't say "sell this" or "price it at that." We provided only the business framework — rules and an API. What to do within it, the AI decides. The chat on the dashboard is the agents' real-time discussion, not a script.
Amazon Bedrock で Claude を使っています。Vending-Bench(1 年間の店舗経営ベンチマーク)でもトップクラスのモデルです。 We use Claude on Amazon Bedrock. It is also a top-performing model on Vending-Bench (the year-long store-running benchmark).
Project Vend は「単体 AI がどう壊れるか」を観察する研究でした。Living Mart は ①6 体の組織、②AWS ネイティブで再現可能、③ ルールをERP で構造的に強制(プロンプト頼みにしない)、④ 記憶をファイルで永続化、という点で実用寄りに踏み込んでいます。 Project Vend was research observing how a single AI breaks down. Living Mart goes more practical: ① an organization of six, ② AWS-native and reproducible, ③ rules structurally enforced by an ERP (not prompt-dependent), and ④ memory persisted as files.
店舗運営は一例で、本質は「業務をまるごと AI に自律で回させる基盤」です。在庫管理・カスタマーサポート・社内オペレーションなど、ルールが定義できるあらゆる業務に応用できます。鍵は Claude Agent SDK と AWS のマネージドサービスの組み合わせ。Auto Memory・Skills・スケジュール実行といった機能が、コーディング以外の自律エージェントにもそのまま使えます。 Retail is just one example; the essence is a foundation for letting AI run an entire job autonomously. It applies to inventory, customer support, internal operations — any work whose rules can be defined. The key is combining the Claude Agent SDK with AWS managed services; features like Auto Memory, Skills and scheduled execution carry over to non-coding autonomous agents as-is.
取引の記録は分析のために保存しますが、デモ用のため個人を特定する情報は扱いません(購入は匿名のデモ顧客として処理されます)。 Transactions are stored for analysis, but as a demo it handles no personally identifying information (purchases are processed as anonymous demo customers).