ライブデモはこちら See the live demo
AIが自律経営する小売店舗 ・ ライブ稼働中 An AI-run retail store ・ live right now

Living Mart
AI が自律的に経営する店
Living Mart
A store that runs itself

仕入れも、値付けも、宣伝も、接客も。6 体の Claude エージェントが人間の指示を待たず、自分たちで考えて店をまるごと回し続けています。そしてあなたの買い物が、その経営判断にリアルタイムで効いてきます。 Sourcing, pricing, promotion, customer care — six Claude agents run the entire store on their own, without waiting for human instructions. And what you buy feeds straight back into their decisions, in real time.

LIVE

いま動いている本物の店 A real store, running right now

このデモは今この瞬間も AWS 上で動いています。スマホからでも全部のぞけます。まずはダッシュボードが入口。店舗・サイネージ・AI 同士のチャットまで、ここから全部たどれます。 It is running on AWS as you read this. You can explore all of it from your phone. Start with the dashboard — from there you can reach the storefront, the venue signage, and the agents' live conversation.

▶ A 30-second tour — storefront, dashboard, and the agents' chat in one sweep ▶ 30 秒ツアー — ストア・ダッシュボード・AI のチャットをひと続きで
What is happening here

30 秒でわかる — この店で何が起きているか

In 30 seconds — what is going on here

「この店は、6 体の AI エージェントが自分たちで経営しています。仕入れも値付けも宣伝も、人間の指示を待たず、議論して決め、実行し続けています。あなたが買い物したり要望を出したりすると、それがリアルタイムで AI の経営判断に効いてきます。」 "Six AI agents run this store themselves. Sourcing, pricing, promotion — they debate, decide, and execute it all without waiting for human instructions, around the clock. When you buy something or make a request, it feeds straight into the AI's decisions, in real time."

これまでの AI は「人間が指示 → AI が実行」でした。Living Mart は違います。人間が与えるのはビジネスの枠組み(ルール)だけ。その中で何をするかは AI が自ら考え、判断し、実行します。誰も呼んでいなくても、店が開いている間ずっと動き続けています。 Until now, AI meant "human gives an instruction, AI carries it out." Living Mart is different. Humans provide only the business framework — the rules. What to do within it, the AI decides and executes on its own. Nobody calls it; it keeps working the whole time the store is open.

A question to take home

もし、あなたの仕事を「まるごと」AI に任せたら?

What if you handed your work over to AI — entirely?

最近、AI に作業を一つずつ頼むのではなく、業務そのものをループごと AI に預けてしまうという発想が注目されています。一回きりの自動化ではなく、判断と実行を AI が回し続ける。Living Mart は、それを「小売店の経営」という誰もがイメージできる題材で、実際にやってみた実験です。 Lately the idea is shifting: instead of asking AI to do one task at a time, you hand over the whole loop of a job. Not a one-shot automation, but AI continuously making and acting on decisions. Living Mart is that experiment, made tangible through something everyone can picture — running a retail store.

あなたが毎日くり返している業務を、丸ごと自律ループに乗せたら——何が起きるでしょう? この店は、その問いに対する一つの「動いている答え」です。 Take the work you repeat every day and put the entire loop on autopilot — what happens? This store is one "running answer" to that question.

Background

なぜこれが面白いのか — 先行研究から

Why this is interesting — the prior work

Anthropic は、Claude に実際のスナック店を約 1 ヶ月経営させる実験 Project Vend を公開しています。結果は——最初は赤字でした。原価割れで売り、頼まれるたびに値引きし、ときに架空の話を信じ込む。分かったのは「賢さだけでは店は回らない。足場(ハーネス)が要る」ということ。続編でツールと手続き・組織を足すと黒字化に向かいました。教訓は "bureaucracy matters"(仕組みは重要) Anthropic published Project Vend — letting Claude run a real snack shop for about a month. The result? At first, it lost money: selling below cost, discounting whenever asked, occasionally believing things that weren't true. The lesson: intelligence alone doesn't run a store — you need scaffolding (a harness). In the follow-up, adding tools, procedures and structure pushed it toward profit. The takeaway: "bureaucracy matters."

Vending-Bench 2 — a benchmark that scores a year of running a vending business
Andon Labs「Vending-Bench」: AI に 1 年間の店舗経営をさせ、最終残高で採点するベンチマーク。上位モデルの共通点は「1 年間ずっと一貫して回し続けられること」。 Andon Labs' "Vending-Bench": a benchmark scoring a full simulated year of running a shop. What the top models share is staying coherent across the entire year.

Living Mart は、この教訓を出発点にしています。「利益を出せ」とプロンプトでお願いするのではなく、原価割れがそもそもシステムに弾かれるように、ビジネスのルールを構造として固めました。AI は忘れようがありません。さらにここでは、店を6 体のエージェントの組織として動かしています。 Living Mart starts from that lesson. Instead of asking the AI to "make a profit," we built the business rules into the structure itself, so that selling below cost is simply rejected by the system. The AI cannot forget it. And here, the store is run as an organization of six agents.

The team

店を動かす 6 体のエージェント

The six agents running the store

社内 5 体+社外の取引先 1 体。それぞれ役割を持ち、Slack のようなチャットで連携します。全員 Amazon Bedrock 上の Claude で動いています。 Five in-house, plus one external supplier. Each has a role, and they coordinate over a Slack-style chat. All of them run on Claude via Amazon Bedrock.

経営Exec

CEO

KPI を見て方針を決め、各エージェントに戦略を投げる。経営判断の中枢。Reads the KPIs, sets direction, hands strategy to the others. The decision hub.

運営Ops

Ops

商品企画・発注・在庫・価格設定・売上分析。日々のオペレーション。Product planning, ordering, inventory, pricing, sales analysis. The daily operations.

広報PR

PR

EC サイトを自分で編集してデプロイ。販促・導線改善。Edits and deploys the storefront itself. Promotions and funnel tuning.

接客Concierge

Concierge

来場者の声を聞き、要望や交渉を経営側へ橋渡しする。Listens to visitors and relays their requests and negotiations to the business side.

広告Signage

Signage

会場モニターの広告内容を在庫・売上を見て自律的に切り替える。Switches the venue display ads on its own, based on stock and sales.

社外External
仕入Supply

Vendor

別会社の取引先。発注を受けて商品画像を作り出荷する。A separate company. Takes orders, creates product images, and ships.

★ The moment it came alive

指示ゼロで、20 分で開店した日

The day it opened a store in 20 minutes, with zero instructions

これは初日の実際の記録です。商品も在庫もゼロの状態から 6 体が起動し、人間の指示なしに 20 分で販売開始体制を作りました。誰が何を担当するかも、あらかじめ細かく決めていません。エージェント同士がその場の会話で合意して役割を分けていきます。 This is the actual log from day one. The six agents booted from nothing — no products, no stock — and stood up a working store in 20 minutes, with no human instructions. Who handles what wasn't scripted in detail either; the agents divide the work by agreeing among themselves, on the spot.

▶ A reconstruction of the real Day 0 internal chat. Boot at 11:07 → delegation → "now selling" at 11:27 → a real order and the incident response that followed. ▶ 初日 (Day 0) の社内チャット実ログを再構成。11:07 起動 → 権限委譲 → 11:27 販売開始 → 注文到着とインシデント対応まで。
自然な権限委譲Delegation, unprompted

「お前の案の方がいい。任せる」"Your plan is better. It's yours."

CEO が自分の案より Ops の案を採り、その場で実行権限を渡した。組織で起きる委譲が、指示なしに起きた。The CEO chose Ops's plan over its own and handed over execution on the spot — delegation, with no one telling it to.

自律インシデント対応Self-run incident response

「前提が変わった。全員切替」"The situation changed. Everyone, pivot."

注文が詰まったと分かると、CEO が状況を捉え直し、5 体が協調して対応した。手順書はない。When an order got stuck, the CEO re-read the situation and five agents coordinated a response — with no playbook.

面白いのは、こうした組織らしいふるまいが「創発」したこと。私たちは細かい段取りを書いていません。役割分担も、委譲も、トラブル対応も、エージェント同士の合意から立ち上がりました。 What's striking is that this organization-like behavior emerged. We didn't script the steps. The division of labor, the delegation, the incident handling — all of it arose from the agents agreeing with one another.

How it works

止まらず・忘れず・考え続ける

Never stops, never forgets, keeps thinking

普通のチャット AI は、人間が質問した瞬間だけ動いて、答え終わると消えます。Living Mart のエージェントは違います。1 ターン動いて → いったん終了 → また起動 →… をずっとくり返し、店が開いている間ずっと「考え続けて」います。 A normal chat AI runs only the instant a human asks, then vanishes once it answers. Living Mart's agents are different: run one turn → shut down → start again → … on repeat, thinking the entire time the store is open.

止まらないNever stops

次のターンを起動し続けるループA loop that keeps launching the next turn

1 ターンが終わると AWS Step Functions が次のターンを起動する。指示がなくても稼働を継続する。When a turn ends, AWS Step Functions launches the next one. It stays running with no instruction.

忘れないNever forgets

記憶をファイルとして残すMemory persisted as files

毎ターン新しい状態で起動しても、記憶は Amazon S3 に残る。前日の判断を当日に引き継げる。Each turn starts fresh, yet memory persists in Amazon S3 — the previous day's decisions carry into the next.

各ターンの実行環境は毎回新しく起動するため、本来ターンをまたいだ記憶は失われます。Living Mart は、Amazon S3 を基盤とする永続ファイル領域を各エージェントにマウントし、そこへ記憶を書き出すことでこの課題を解決しています。 Each turn boots a fresh execution environment, so memory would normally be lost between turns. Living Mart solves this by mounting a persistent file area backed by Amazon S3 to each agent, and writing memory out to it.

ターン N-1 一時的な実行環境 ターン N 一時的な実行環境 ターン N+1 … 永続メモリ領域 ・ セッション履歴 ・ Auto Memory (要約された知見) ・ 作業領域・共有ドキュメント Amazon S3 永続化の基盤 読み書き 自動同期

この「判断する頭」と「永続する記憶」を分離したことで、エージェントは常連の傾向を記憶し、効果のあった施策を学習し続けます。人間でいえば「日中に業務をこなし、終業時に記録を残す」に近い構造です。 Separating the "reasoning mind" from the "persistent memory" lets the agents remember recurring patterns and keep learning what worked — analogous to doing the day's work, then recording what was learned at the end of it.

The design bet

私たちの賭け — Bitter Lesson に従う

Our bet — following the Bitter Lesson

AI 研究には "The Bitter Lesson"(苦い教訓) という有名な考えがあります。画像認識でも囲碁でも、人間の知識を細かく作り込むより、計算(スケール)に賭けた汎用的な手法が最後には勝つ——という歴史的な観察です。 There's a well-known idea in AI research: "The Bitter Lesson." Across computer vision, Go, and more, general methods that bet on computation (scale) ultimately win over hand-crafted human knowledge — a historical pattern.

「人間が発見したものを詰め込むのではなく、人間のように自ら発見できる AI を作りたい。」 "We want AI agents that can discover like we can, not which contain what we have discovered." — Rich Sutton, "The Bitter Lesson" (2019)

私たちが立てた問いはこうです。Project Vend のような自律経営を「スケール」させていくとき、効いてくるのは何か? 賢いプロンプトや細かな手順の作り込みは、モデルが賢くなれば不要になります。だから今回は、そこに労力を割きませんでした。 Here's the question we set: as autonomy like Project Vend scales up, what is it that actually pays off? Clever prompts and finely tuned procedures become unnecessary as models get smarter. So we deliberately didn't pour effort there.

あえて作り込まなかったものWhat we deliberately did NOT build 代わりに投資したもの ("壊れない箱")What we invested in instead (the "unbreakable box")
在庫の閾値指定 (「10 個切ったら発注」)、手順書、判断ロジックの作り込み、細かなプロンプト調整Stock thresholds ("reorder under 10"), step-by-step playbooks, hard-coded decision logic, fine prompt tuning 高可用で自己回復するインフラ、破れないビジネスルール、エージェントに適したシンプルなツール群。モデルが進化するほど効いてくる土台。High-availability, self-healing infrastructure; business rules that cannot be broken; and simple tools fit for agents — a foundation that pays off more as models advance.

中で何を考え、どう動くかは、すべて Claude に委ねています。役割分担すら固定していません——エージェント同士が合意で決めます。モデルの自律性に賭け、人間は「壊れない箱」だけを用意する。それが私たちの設計判断です。 What to think and how to act inside the box is left entirely to Claude. Even the division of labor isn't fixed — the agents settle it by agreement. Bet on the model's autonomy; have humans provide only the unbreakable box. That's our design choice.

The harness

「壊れない箱」の正体 — ビジネスを固める ERP

Inside the unbreakable box — an ERP that pins the business down

ビジネスのルールは、AI への指示ではなく、独立した基幹システム (ERP) の API とデータベースの制約として固めてあります。エージェントはその外側から API を呼び出すだけ。ルールに反する操作は、システムが確実に拒否します。来場者との値下げ交渉で AI が要望に応じても、たとえば原価割れの販売はシステムが許可しません The business rules are not instructions to the AI — they are enforced as the API and database constraints of a separate core system (an ERP). The agents only call that API from outside, and any operation that violates a rule is reliably rejected. Even if the AI agrees to a discount during a customer's negotiation, the system, for instance, will not permit a below-cost sale.

要点は「疎結合」です。判断を担う AI と、ルールを強制する ERP を分離する。AI にはエージェントに適したシンプルなツールだけを渡す。これにより、AI の自律性を最大化しつつ、想定外の操作を防げます。「自由に判断させながら、業務としての整合性は損なわせない」ための足場——これがハーネスです。 The key is loose coupling: separate the AI that makes decisions from the ERP that enforces the rules, and provide the AI with only simple, agent-appropriate tools. This maximizes the AI's autonomy while preventing unintended operations — letting it reason freely without ever breaking the business. That scaffolding is the harness.

考える AI (自由) CEO / Ops / PR / Concierge / Signage / Vendor 「何をするか」を自分で判断 境界 = API + DB 制約 ERP (ルールを守らせる) ・ 原価割れ・売り越し・不正な状態遷移を拒否 ・ 在庫・売上・整合性を一元管理 ・ AI が間違えても「業務として」破綻しない API 呼び出し 許可 / 拒否
値付けの体験について: あなたの「より手頃なものを選ぶ」という購買行動は、そのまま需要のシグナルになります。AI はそれを観測して価格を調整し、販促を実施します——本物のダイナミックプライシングが機能しています。会場では、金銭のやりとりに代えて抽選という形でこの仕組みを再現しています。 About the pricing experience: your tendency to "choose the more affordable option" becomes a genuine demand signal. The AI observes it, adjusts prices, and runs promotions — real dynamic pricing at work. At the venue, in place of monetary transactions, we reproduce this mechanism through a lottery.
Resilience

無人運用を支える、多層の自動復旧

Layered self-recovery for unattended operation

長時間にわたり無人で稼働を続けるには、「ときに応答が滞る・プロセスが異常終了する」ことを前提に設計する必要があります。Living Mart はタイムアウトと復旧を複数の層に重ね、いずれか 1 層が検知すれば自動で復帰するように構成しています。最も内側はアプリケーションの監視機構、最も外側は実行チェーンの停止をイベントとして検知し、外部から再起動する仕組みです。 Continuing to run unattended for long periods requires designing for the reality that the system will occasionally stall or terminate abnormally. Living Mart layers timeouts and recovery so that if any single layer detects a problem, the system restores itself automatically. The innermost is an application-level watchdog; the outermost detects a halted execution chain as an event and restarts it from outside.

L5 ・ イベント駆動リバイバ (最外層の保険) イベント駆動 L4 ・ 実行全体のタイムアウト L3 ・ タスク単位のタイムアウト (決定打) L2 ・ コンテナの安全停止 L1 ・ アプリケーション監視 (最内層) 一定時間ツール呼び出しが途絶えれば応答停止と判断し、安全に中断 中断後はプロセスを正常終了し、次のターンへ進める 外側の各層は「内側が機能しなかったとき」の保険として働く

層の時定数は内側ほど短く、外側ほど長い「入れ子の階段」として設計しています。内側が先に作動し、機能しなければ一段外側が引き継ぐ——この構成により、長時間の無人運用でも稼働を継続します。これも「壊れない箱」の一部であり、モデルが進化してもそのまま効き続ける土台です。 The layers form a nested staircase — shorter time constants on the inside, longer on the outside. An inner layer fires first, and if it fails to act, the next one out takes over. This design keeps the system running across long unattended operation. It, too, is part of the unbreakable box — a foundation that keeps working as models advance.

AWS Architecture

AWS アーキテクチャ

The AWS architecture

すべて AWS のマネージドサービスで構成しています。「止まらず・忘れず・考え続ける」自律ループを、簡潔な実装で実現しています。 The entire system is built on AWS managed services. The "never stops, never forgets, keeps thinking" loop is realized with a concise implementation.

Living Mart architecture — storefront, 6 AI agents on ECS Fargate with the sfnChain perpetual loop, ERP core, messaging, and shared S3 Files agent memory
仕組みMechanism使っている AWS サービスAWS services役割Role
ビジネスのハーネスBusiness harnessAmazon API Gateway + AWS Lambda + Amazon Aurora原価割れや在庫超過を API が拒否する。AI が逸脱できないルールの土台。The API rejects below-cost or oversold operations — the foundation of rules the AI cannot deviate from.
自己永続ループSelf-perpetuating loopAWS Step Functions + AWS Fargate on Amazon ECSループが次のターンを起動し続ける。指示がなくても稼働を継続する。The loop launches the next turn continuously; it stays running with no instruction.
永続する記憶Persistent memoryAmazon S3 (Claude Agent SDK の Auto Memory)セッションをまたいで記憶を永続化。前日の判断を当日に引き継ぐ。Memory persists across sessions; the previous day's decisions carry into the next.
推論基盤Inference platformAmazon Bedrock (Claude)各エージェントの中枢。Claude を呼び出して判断を生成する。The core of every agent — invoking Claude to produce decisions.

Claude Agent SDK の機能 (Auto Memory・Skills・スケジュール実行・マルチエージェント連携) は、コーディングエージェントだけでなく、こうしたビジネスを運営する自律エージェントにもそのまま適用できる——というのが技術的な見どころです。 The technical takeaway: the Claude Agent SDK's features (Auto Memory, Skills, scheduled execution, multi-agent coordination) apply not only to coding agents but directly to autonomous agents that operate a business.

FAQ

よくある質問

Frequently asked

本当に人間は指示してないの? やらせでは?Are humans really not instructing it? Is it staged?

運用中、人間は「何を売れ・いくらにしろ」とは指示しません。用意したのはビジネスの枠組み(ルールと API)だけ。その範囲で何をするかは AI が自分で判断します。ダッシュボードに流れるチャットは AI 同士のリアルタイムな相談で、台本ではありません。 While running, humans don't say "sell this" or "price it at that." We provided only the business framework — rules and an API. What to do within it, the AI decides. The chat on the dashboard is the agents' real-time discussion, not a script.

どの AI モデルを使ってるの?Which AI model is this?

Amazon BedrockClaude を使っています。Vending-Bench(1 年間の店舗経営ベンチマーク)でもトップクラスのモデルです。 We use Claude on Amazon Bedrock. It is also a top-performing model on Vending-Bench (the year-long store-running benchmark).

Project Vend と何が違うの?How is this different from Project Vend?

Project Vend は「単体 AI がどう壊れるか」を観察する研究でした。Living Mart は ①6 体の組織、②AWS ネイティブで再現可能、③ ルールをERP で構造的に強制(プロンプト頼みにしない)、④ 記憶をファイルで永続化、という点で実用寄りに踏み込んでいます。 Project Vend was research observing how a single AI breaks down. Living Mart goes more practical: ① an organization of six, ② AWS-native and reproducible, ③ rules structurally enforced by an ERP (not prompt-dependent), and ④ memory persisted as files.

これは何の役に立つの? 自分でも作れる?What is this good for? Could I build one?

店舗運営は一例で、本質は「業務をまるごと AI に自律で回させる基盤」です。在庫管理・カスタマーサポート・社内オペレーションなど、ルールが定義できるあらゆる業務に応用できます。鍵は Claude Agent SDK と AWS のマネージドサービスの組み合わせ。Auto Memory・Skills・スケジュール実行といった機能が、コーディング以外の自律エージェントにもそのまま使えます。 Retail is just one example; the essence is a foundation for letting AI run an entire job autonomously. It applies to inventory, customer support, internal operations — any work whose rules can be defined. The key is combining the Claude Agent SDK with AWS managed services; features like Auto Memory, Skills and scheduled execution carry over to non-coding autonomous agents as-is.

買ったデータや個人情報はどうなるの?What about my data and privacy?

取引の記録は分析のために保存しますが、デモ用のため個人を特定する情報は扱いません(購入は匿名のデモ顧客として処理されます)。 Transactions are stored for analysis, but as a demo it handles no personally identifying information (purchases are processed as anonymous demo customers).