How We Built a Real Estate Platform for 150,000 Listings
Last December we started working on the idea of a new real estate platform for the Greek market: a place where people search for homes to buy or rent, and agencies publish their listings straight from their own systems. We had 130 properties in the database. By July HouseMaster had about 85,000, and on the last weekend of September our weekly duplicate run went through 130,209 of them. It now holds 150,000.
I've been the Technical Lead of the team since that December. Most of the engineering went into 3 problems: search, duplicates and ingestion. What follows is what each one turned out to be, what we shipped for it, and the decisions that held up along the way.

Everything in this post is the team's work. I've worked closely throughout with our CEO, Marietta L., to define the product and the technical direction, and with our support specialist, Christina L., to onboard hundreds of real estate agencies around Greece. Dimitris K. set up and maintains our DevOps and the Kubernetes cluster all of it runs on, and Manos N. has been responsible for the application layer.
Here is what this post covers:
- What runs where: the architecture, the cluster, and what sits outside it.
- Search: how the platform's core feature actually works, from the map query to the order of the results.
- Duplicates: finding the same flat among 150,000 listings every night, without affecting user traffic.
- Ingestion: taking listings from every agency's CRM and keeping the data clean.
- What we learned: the decisions that held up, from keeping geography in the database to isolating the nightly work.
What Runs Where
The frontend Manos built is a Next.js app, and the backend is a FastAPI API. Both run on AWS, in Kubernetes on EKS, each as its own Helm release. A merge to main in any of the 3 repositories builds the image on GitHub Actions, pushes it to ECR and runs helm upgrade.
The shape is 3 services in one cluster, and it buys us 3 things. Each service deploys on its own, so a frontend change never redeploys the API, and a bad release rolls back alone. The API scales on CPU and memory while the nightly work runs somewhere else: the ingestion service is a second release, a small internal API that accepts jobs and an ARQ worker that processes photos, enriches locations and finds duplicates at night, so none of it competes with user traffic. And the scheduled jobs, notifications, ranking scores and price snapshots, are CronJobs on the API's own image, so they share its code without being a fourth thing to deploy.
Everything stateful lives outside the cluster, which is what makes a pod disposable: PostgreSQL with PostGIS for everything geographic, Redis for caching, the ranking scores and the job queue, and S3 behind CloudFront for the property photos. Sign-in goes through the API to Supabase Auth, payments to Stripe, email to SendGrid, and Logfire traces every request and query.
That is the platform. The first problem it had to solve is the one every visitor meets within seconds of arriving.
Searching 150,000 Listings on a Map
Search is what people come to HouseMaster for. A search can be scoped in 3 ways: a map viewport, a radius around a point, or a set of areas from the location tree. The map is the one people use most, and it was the first thing to break at scale.

Every move of the map asks the database for the listings inside the visible area. Doing that comparison on a globe is the accurate way, but as the dataset grew it became the slow one: very wide map views came back empty or failed outright. So the map query now uses a cheaper comparison with an index built for it, and that alone cut its p95 by more than half.
What a Search Has to Do
Whichever way a search is scoped, it has to do 4 things:
- Return only the listings that match the filters inside that scope.
- Put the newest listings first.
- Not let a single agency fill a page of results.
- Answer quickly on the first pages, whatever the filters.
The map query above takes care of the first. The other three are about order, and they pulled against each other.
The Ranked Window
The bigger problem was order. Sorting purely by newest meant an agency that had just pushed a batch could fill the first pages on its own, and a visitor would see one agency's stock before anything else. So every search now builds a ranked window first:

- We take a fixed window of the newest listings that match the filters.
- Each one gets a score. Ranking considers the listing's quality, its freshness, how people respond to it, its price against the area and any recent price drop.
- The scored list is reordered so that no single agency fills a page of results.
- The result is cached in Redis for a short while, under a key that ignores the page number.
Pages inside the window are just slices of that cached list, so paging through the first results doesn't run the query again. Beyond the window, results go back to newest first.
The diversity step took a few tries. A simple greedy pass gets stuck when 2 big agencies dominate the window, because by the end only their listings are left and every slot breaks a rule. So we tried several variants and checked each against an exhaustive search over 1,481 small inputs. The version we shipped found a valid order every time one existed, while the variant we dropped failed on 37 of them.
💡 Rank a fixed window once and cache it, so the expensive work happens on the first page only.
The response scores are the costly part of that ranking, and caching them cut the p95 of search by more than 40% after the release, while the share of requests over 2 seconds halved. Keep in mind that this compares the first hours after the release with the hours before it, so a full-day number is still owed.
The Same Flat, Listed by Five Agencies
A problem most real estate platforms face is the same flat listed by several agencies, each with its own copy of the photos, a slightly different address and its own price. Showing it 5 times in a row makes the whole catalogue feel padded, so the ingestion service groups those listings every night.

Comparing 150,000 listings with each other would mean billions of pairs, so we only compare listings that fall in the same geographic bucket. Matching then considers each pair's address, attributes like size, rooms and floor, the photos and the location, and the pairs that match are joined into groups.
Making the Nightly Run Finish
The nightly run is the ingestion service's duplicate job. Each night it takes the listings that changed, compares them with the ones near them and joins the matches into groups, and once a week a full run does the same for the whole catalogue.
A full run compares millions of pairs, and the first versions of it didn't get to the end. One hit the database's statement timeout, which splitting the geographic query into smaller chunks fixed. The next ran out of memory, and raising the pod's limit only moved the failure a little later.
tracemalloc, Python's own tool for seeing which lines of code the memory belongs to, showed why. To avoid scoring the same 2 listings twice, the job wrote down every pair it had already looked at, and it kept that list in one place for the whole run. Each entry costs about 250 bytes, which sounds like nothing, but the run had around 15.6 million pairs to look at, and 15.6 million entries at 250 bytes each is close to 4 GiB. The list of what the job had already done was bigger than the pod it ran in, and no memory limit we could reasonably set would have held it.

The fix was to process the changed listings bucket by bucket and keep the seen-pairs set per batch, so it's thrown away after each one. A full run now goes through 15.6 million pairs in about 2 minutes. A nightly incremental run takes about 10 minutes and the weekly full run about 15.
💡 When a job gets OOMKilled, measure what's growing before raising the limit.
Getting the Matches Right
A precision check found grouped pairs whose 2 listings were far apart. Almost all of them matched on address alone, because some agencies stamp their office address on every parking space and plot they publish. So distance now counts against a match, whatever else agrees, and the far-apart pairs dropped by more than half.
We had also missed something simpler. The job that prepares the photos for comparison existed, but it had never been scheduled, so only a few percent of the listings had been through it. Once it ran every night, coverage climbed within days.
Listings Pushed From Agency CRMs
Agencies' CRM systems push their listings to us through a batch endpoint, each agency with its own key, and every agency's CRM has its own idea of what a listing is. Any batch can hold 1 record that is wrong in a new way, so each property in a batch is written in its own database savepoint, and one bad record doesn't fail the rest. The response carries per-item errors plus data-quality warnings, for example a property marked as new construction whose build year is decades old.

Every listing is then matched to a three-level location tree, prefecture, neighbourhood and about 15,000 sublocations, using neighbourhood boundaries where they exist and the nearest centre otherwise, with a confidence score that feeds an admin report. The public map shows a slightly offset position, not the exact address. The ingestion worker enriches each area with nearby points of interest from Google Places on a fixed nightly budget.
Put together, this is how we take in thousands of properties a day from systems we don't control. Each agency pushes on its own key, every record is accepted or rejected on its own, and every accepted one lands in the location tree with a warning attached if its data looks off.
What We Learned
Search, duplicates and ingestion broke in different places and at different sizes, but the fixes kept coming back to the same few decisions. Ten months in, with the catalogue a thousand times bigger than the one we started with, these are the ones that have carried the most weight. Some were made on day 1 and have never needed revisiting; others came out of a job that wouldn't finish or a query that wouldn't return. I would make them again on the next platform of this shape.
Keep geography in the database. Every map, radius and area query runs in PostGIS next to the listing data. A separate search engine would have been faster to bolt on, and it could have drifted from the listings every time a sync failed.
Rank a window, not the table. Scoring every match on every page is what makes search slow, and sorting by newest is what lets one agency own the first pages. Ranking a fixed window once, diversifying it and caching it fixed both at the same time.
Compare less before comparing faster. Finding duplicates among 150,000 listings is a blocking problem before it is a matching problem. Deciding which pairs never to compare, and keeping the seen-pairs set per batch, mattered more than anything in the matching itself.
Isolate the nightly work. Photos, enrichment and duplicates run in their own release and their own worker, so the API never feels them, while the API's own scheduled jobs ride on its image and need nothing extra to deploy.
Accept imperfect data at the edge, and say what was wrong. A bad record from a CRM fails on its own, never the batch, and the response names the problem. That is cheaper than cleaning data after it is in.
Measure before you resize. Raising a memory limit moved the failure later; measuring what was growing removed it. The same goes for a slow query and an index.
Make the repository the truth for what runs. Runtime settings in the chart, one command that runs every test, migrations that give way to readers, and the live image tag read after every deploy. None of this is specific to real estate, and all of it came from the gap between the repository and production.
Those decisions are what carried the platform from 130 listings to 150,000, and the catalogue is still growing, along with the list of what we want to build on top of it. You can see the platform at housemaster.gr, there's a short summary of the project in my portfolio, and Building a Website in 2026 has more details on the stack behind it.