Lakehouse Prepchief architect

#The Full Mock Loop

Run this on Day 10. Timed, out loud, recorded. Ideally with a person playing the panel β€” pay a friend in dinner if necessary; a live human who interrupts is worth more than another day of reading.

Total: ~3.5 hours. Do it in one sitting, because stamina is part of what is being tested.


#Round 1 β€” Recruiter screen (15 min)

Answer these in order, aloud, without notes:

  1. Walk me through your background. (90 seconds. Time it.)
  2. Why are you looking?
  3. Why this role specifically?
  4. What are you looking for in your next move?
  5. Talk me through your current scope β€” team, budget, decisions you own.
  6. Compensation expectations?

Scoring

1 3 5
Arc Rambling chronology Clear but long 90 seconds, ends at why-this-role
Motivation Generic Plausible Specific to this company's actual problem
Scope Vague Described Quantified β€” people, budget, decisions

#Round 2 β€” Hiring manager (45 min)

  1. Tell me about the most complex Databricks environment you've architected.
  2. How do you approach a customer/engagement you know nothing about in week one?
  3. Tell me about a time you disagreed with a customer's technical direction.
  4. What's the worst architecture decision you've made?
  5. How do you stay current? Give me something that changed in the last six months and why it matters.
  6. How do you handle being the most senior technical person in a room where you're wrong?
  7. What would your last team say is your biggest weakness?

Scoring

1 3 5
Specificity No examples Examples without numbers Named constraints, numbers, decisions you personally made
Ownership "We" throughout Mixed Clear "I decided", team credited separately
Self-awareness Only successes One soft failure Real failure, real learning, applied since
Currency (Q5) Nothing recent A feature name A change and its architectural consequence

#Round 3 β€” Technical deep dive (60 min)

Run these end to end. No looking anything up. Each should take 4–8 minutes.

  1. Walk me through what happens when a SQL query runs on Databricks, from submission to results.
  2. What's in the Delta transaction log, and how does it give you ACID on object storage?
  3. Two writers commit to the same table at the same moment. What happens?
  4. A job that took 20 minutes now takes 70. Nothing changed. Diagnose it.
  5. One task in a stage runs 40 minutes; the other 199 finish in 30 seconds. What is it and what do you do, in order?
  6. Explain AQE. What can it not fix?
  7. A customer's MERGE takes four hours. Work through it.
  8. When is Photon not worth it?
  9. Partitioning, Z-order and liquid clustering β€” pick one for a 50 TB event table and defend it.
  10. How many Unity Catalog metastores for an organisation in the EU and US with dev/test/prod?
  11. A CISO says serverless is unacceptable. Respond.
  12. What must be true for end-to-end exactly-once in Structured Streaming?
  13. Their RAG system gives bad answers. Where do you start and why not the prompt?
  14. When would you tell a customer not to use Databricks?

Scoring

1 3 5
Accuracy Material errors Broadly right Precise, including the limits
Depth on demand One layer Two Three+, and knows where its own knowledge ends
Diagnostic method Jumps to fixes Some structure Evidence before hypothesis, every time
Honesty Bluffs Hedges Says "I don't know" cleanly, then reasons well

Mark any question where you bluffed. That is the highest-priority gap, above any knowledge gap.


#Round 4 β€” Panel presentation (60 min)

The brief (give yourself 45 minutes to prepare, no more β€” this simulates the real constraint):

FinServ Group β€” a European insurance and wealth business, Β£4bn revenue, operating in the UK, Germany and Ireland.

They run a 12-year-old Teradata warehouse (contract renewal in 16 months, quoted at +45%), an on-prem Hadoop cluster used mainly by the actuarial team, and a sprawl of SQL Server marts. Around 40 data engineers and analysts across three business units that do not get on.

The regulator has issued a finding: they cannot evidence who accesses customer PII. The CEO has announced an "AI-first" strategy to the market and expects something visible this year. The CFO has told everyone that the data budget is not increasing.

You have 30 minutes with their CIO, Head of Data, CISO and a lead engineer. Present your approach.

Build the deck per ../02-drills/05-the-panel-presentation.md. Present for 30 minutes. Have your panel interrupt at least six times, including:

  • One hostile: "We've heard this pitch before and the last programme failed."
  • One CFO: "What does this cost and when do we see a return?"
  • One CISO: "Where exactly does our data go?"
  • One engineer: "Why not just put Iceberg on S3 ourselves?"
  • One off-topic, to test whether you return to the thread.
  • One deliberately flat non-reaction, to test whether you keep energy without feedback.

Scoring β€” use the table in ../02-drills/05-the-panel-presentation.md Β§1. Score honestly on discovery and time management first; they predict the rest.

Note the trap in this brief: three stakeholders with conflicting priorities (regulator, CEO's AI mandate, CFO's flat budget) and a hard contract date. A strong answer sequences so that the governance work closing the regulatory finding is also the foundation for the AI mandate, and pays for itself by unlocking the Teradata decommission. If your answer treats them as three separate programmes, you have missed the exercise.


#Round 5 β€” Values / behavioural (30 min)

One question per value, three minutes each, no notes:

  1. Customer obsessed β€” a time you advocated for the customer against internal pressure.
  2. Raise the bar β€” a time you raised the standard of work around you.
  3. Truth seeking β€” a time you were wrong, and what it cost.
  4. First principles β€” a time you solved something by going back to fundamentals.
  5. Bias for action β€” a time you moved without complete information.
  6. Company first β€” a time you gave something up for a better overall outcome.

Scoring

1 3 5
Structure Meandering STAR present STAR-L, 3 minutes, no notes
Ownership "We" Mixed "I decided…", team credited
Result None Qualitative A number
Learning Absent Generic Specific, and visibly applied since

#Round 6 β€” Executive panel (20 min)

Answer as if to a CFO and a CIO:

  1. Sell me this platform investment in five sentences.
  2. What's the payback period, and what's the assumption most likely to be wrong?
  3. What's the cheapest version that still closes the regulatory finding?
  4. Why now?
  5. Our last data programme failed. Why is this different?
  6. What have you left out of your proposal?

Scoring: was there a number in every answer? Did you volunteer the parallel-running cost before being asked? Did you name a real omission on Q6?


#Scoring the whole loop

Round Score /5 Biggest gap
Recruiter
Hiring manager
Technical deep dive
Panel presentation
Values
Executive

Readiness thresholds:

  • Any round below 3 β†’ that is your entire afternoon.
  • Presentation below 4 β†’ rehearse it twice more; it carries the most weight in a field loop.
  • All rounds at 4+ β†’ stop studying. Rest is now worth more than revision.

The one question to answer after the mock: where did I bluff? Go fix those first β€” not the things you didn't know, but the things you pretended to know. Interviewers detect the difference, and only one of them is disqualifying.