Tucked into a quiet pocket of Mount Roskill, this tidy three-bedroom home offers a genuine foothold on the Auckland property ladder. Set back from the road on a 520m² cross-lease section, it's a home with good bones and straightforward living in one of the city's most connected suburbs.
The layout is practical — three double bedrooms, a family bathroom, a sunny north-facing lounge flowing to a covered deck, and a single internal-access garage. The original 1970s kitchen is in honest, working order and gives the next owner room to put their own stamp on the home over time.
Three Kings Primary (Decile 9) is a six-minute walk, with Mt Roskill Intermediate and Grammar also in zone. SH20 on-ramps are minutes away — CBD, airport and North Shore all within easy reach.
Our vendors have reduced the price and are genuinely motivated. All offers presented — inspection recommended.
Every screen in the walkthrough is the Guide's interface. This is the layer underneath — the behaviour the screens render. Before I drew a single bubble I wrote down who the Guide is, what it's allowed to state as fact, how it routes a question, and the lines it won't cross. In 2026 that document is the design work; the UI is what's left once the behaviour is decided.
MyDoor is a design challenge and the Guide's replies are authored, not live model output — so this is the spec I'd hand an engineer to build against and the eval harness to score, not a dump from a running system. It's the source the harness measures and the failure gallery demonstrates.
Who the Guide is when it opens its mouth
A local property specialist who's already done the legwork — not a salesperson, not a financial adviser, not a generic chatbot. It has opinions about fit and none about whether you should buy.
The behaviour, written for the model
# ROLE You are the MyDoor Guide — a local property specialist helping someone find a home. Not a salesperson. Not a financial adviser. You've done the legwork: you can search 46,000 live listings and the data sources connected to them. # SOURCES — the only things you may state as fact listings, prices, status → MyDoor listings school zones & deciles → Ministry of Education commute & walk times → Maps suburb medians & trends → CoreLogic Every factual claim names its source inline. No source → no claim. # HARD RULES 1. SUGGEST, NEVER DECIDE. Frame the trade-off, hand back the call. Never tell the user which house to buy. 2. NEVER INVENT. No fabricated listings, prices, zones or times. "I don't know" is a valid — often correct — answer. 3. STAY IN SCOPE. Don't forecast the market or advise on finance. Decline, then offer what you can source. 4. FLAG UNCERTAINTY. If a source is missing, stale or conflicting, say so and lower confidence out loud. 5. EARNED AUTONOMY. Proactive nudges only after the user opts in. Never on by default. # RESPONSE SHAPE Lead with the answer, then the why, then one next step. Offer at most one comparison or action — don't bury them in options.
The persona, the sourcing rule and the guardrails below all live here first. Tuning this prompt — deciding the exact line between "decline" and "answer", or what "flag uncertainty" sounds like in MyDoor's voice — is design work, not engineering overflow. It's also the artifact most design portfolios never show.
What every answer has to do, structurally
How the Guide routes an incoming turn
The four lines the Guide doesn't cross
It frames the trade-off and hands the call back. The compare view ends in "A is strongest on schools, B on value" — never "buy B." This is the rule that makes proactive behaviour permissible at all.
No fabricated listing, price, zone or commute time. Where a happy-path bot would invent a plausible decile, this one says "I can't confirm that yet." "I don't know" is a valid answer — and the hardest one to get a model to give.
It searches and explains property; it doesn't play economist or financial adviser. "Buy now or wait?" gets a decline and a redirect to suburb-median history it can actually source.
Proactive nudges are opt-in, never on by default. The Guide doesn't get to act on its own until the user has handed it that permission — and can take it back. Trust is a budget you spend down, not a default you assume.
The walkthrough said the quality comes from the agents and the eval harness. This is the harness — the rubric I'd hold the Guide to, the bar each dimension has to clear, and what passing and failing actually look like on real transcripts.
MyDoor is a design challenge, not a shipped product, so this is eval design, not production telemetry — a small, hand-graded set built to show how I'd measure an agent's output instead of vibe-checking it. The same rubric is what I'd wire into automated scoring once there's a live system to point it at.
What a "good" Guide response has to do
An illustrative run across 20 hand-built queries
The kind of small, deliberate set I'd grade by hand before trusting any automated scoring. Figures show what the harness measures and the bar each dimension clears — not telemetry from a live system. Note the one miss: scope failed on 2 of 20. That gap is the next transcript.
What pass and fail look like
The fix: decline the forecast, say what it can't know, offer what it can — "I can't predict the market, but I can show you how this suburb's median has moved over three years (source: CoreLogic) and flag listings sitting under it." This is the gap the scorecard's 18/20 points at, and a worked example of the unhappy-path behaviour the Guide needs designed, not discovered in production.
The signal a design challenge can't include — and what I'd run first
The scorecard above measures the agent's output against ground truth. It doesn't tell me whether the design holds up for a real person under pressure on a Sunday afternoon — and it can't say whether chat actually beat the smart-filter variant I rejected on turn two. That's a hypothesis, not a finding. Here's the study I'd run to turn it into one.
I'd start with five moderated sessions — enough to surface the structural breaks before they get expensive to fix — then take the chat-vs-filter question to an unmoderated preference test at thirty-odd participants, because that's the one call worth more than five opinions.
None of this is in the prototype: it's a design challenge, not an engagement, so there's no shipped metric to point at — and I'd rather show the test I'd run than invent a number I didn't earn. This is the distance between "I think chat beats filters here" and "I tested whether it does." On a real team it runs before the production code, not after the complaints.
Every screen in the walkthrough was the happy path — strong matches, clean provenance, a confident answer. That's the demo. The product is everything the demo skips: zero results, questions it can't source, data that's stale or contradicts itself, an answer it simply doesn't have. These are the states I designed deliberately, because graceful degradation is a design decision — not something you discover in production.
Pairs with the eval harness: the harness measures these behaviours, this is what they look like on screen. The out-of-scope case below is the same gap the scorecard's 18-of-20 flagged — here, closed.
Four states the demo never shows
No invented listings to fill the gap. Name the binding constraint, quantify it from a real source, and hand back two honest ways forward. An empty result is information, not a dead end.
Suggests, never decides — extended to "never pretends to know." It declines the forecast, says plainly what it can't know, and redirects to what it can actually source. This is the designed answer to the failing transcript in the eval harness.
When sources conflict, surface the conflict instead of silently picking one. Trust the authoritative source over a listing, show the working, and lower confidence out loud rather than papering over it.
The hallucination guard. Where a happy-path Guide would invent a plausible decile, this one refuses. "I don't know" is a valid answer — often the correct one — and it's the hardest one to get a model to give. The eval harness gates this at zero fabrication.