/writing/recruiting & ats/python-interview-questions-for-data-engineer
§ Hiring Tips·20 min read·October 5, 2026

Python Interview Questions for Data Engineers: 12 to Know

O
Olibr TeamHiring Tips
Python Interview Questions for Data Engineers: 12 to Know

Python Interview Questions for Data Engineers: 12 to Know

Data engineering interviews rarely fail candidates on syntax. They fail them on judgment: how you handle a 20 GB file, a messy schema, or a pipeline that breaks at 3 a.m. If you are searching for python interview questions for data engineer roles, you want the questions that actually get asked, with answers you can say out loud.

This list gives you exactly that. Each of the 12 questions comes with a short, direct answer and a working code example you can adapt. They cover both conceptual topics, such as generators, memory management, and the GIL, and hands-on coding challenges like deduplicating records, parsing nested JSON, and cleaning data with pandas.

We also write this from the hiring side. At Olibr, we help recruiters screen and match data engineering candidates every day, so we see which answers separate strong candidates from average ones. Recruiters can use the same 12 questions as a ready-made screening checklist. Start with the first question below, and practice writing each solution without looking.

1. Explain generators and lazy evaluation for large datasets

Sample answer

A conveyor belt feeding one box at a time from a large warehouse into a workstation.

A generator is a function that uses yield to hand back one value at a time instead of building a full list. Python pauses the function after each yield and resumes it on the next request. This is lazy evaluation: values are computed only when needed, so memory stays flat no matter how large the input is.

In a pipeline, you use generators to stream rows from a file, transform them, and write them out in batches. A list of 50 million rows can eat many gigabytes. A generator holds one row at a time. The tradeoff is that a generator is single-use, and you cannot index it or call len() on it.

Generators let you process data of any size with the memory footprint of a single record.

Code example

Here is a three-stage streaming pipeline that reads a large CSV, filters rows, and groups them into batches for loading.

import csv

def read_rows(path):
    with open(path, newline="") as f:
        for row in csv.DictReader(f):
            yield row

def active_only(rows):
    for row in rows:
        if row["status"] == "active":
            yield row

def batches(rows, size=1000):
    batch = []
    for row in rows:
        batch.append(row)
        if len(batch) == size:
            yield batch
            batch = []
    if batch:
        yield batch

for batch in batches(active_only(read_rows("users.csv"))):
    load_to_warehouse(batch)

Each stage pulls from the one before it, so only one batch lives in memory at any moment. Nothing runs until the for loop starts asking for data.

What interviewers look for

Strong candidates go beyond defining yield. They show they have used generators on real volumes and know where they break down. Listen for these points:

  • The difference between a list comprehension [x for x in data] and a generator expression (x for x in data) in memory use.
  • Awareness that generators are exhausted after one pass, and that you must recreate them to iterate again.
  • Chaining generators the way the example does, and linking the idea to chunked reads in pandas or database cursors.

Watch for one red flag. Candidates who say generators are simply "faster" have missed the point. They are about memory control, not speed, and a good answer says so.

2. Compare lists, tuples, sets, and dictionaries

Sample answer

Each type fits a different access pattern. Lists hold ordered, mutable sequences. Tuples are immutable, so they work as dictionary keys. Sets store unique values with constant-time membership checks, and dictionaries map keys to values with fast average O(1) lookups.

Type Ordered Mutable Duplicates Lookup by value or key
list Yes Yes Allowed O(n) scan
tuple Yes No Allowed O(n) scan
set No Yes Not allowed O(1) average
dict Insertion order Yes Unique keys O(1) average

Choose a data structure by how you will access the data, not by habit.

Code example

Checking x in my_list scans every item, while a set jumps straight to it. This snippet filters new records against known IDs and builds a dictionary index for joins.

seen_ids = {r["id"] for r in old_rows}          # set
new_rows = [r for r in rows if r["id"] not in seen_ids]

by_id = {r["id"]: r for r in rows}              # dict
key = (r["id"], r["region"])                    # tuple as composite key

What interviewers look for

Strong candidates explain time complexity, not just definitions. They tie each structure to a pipeline task, such as deduplication or lookups. Listen for these points:

  • Why a list cannot be a dictionary key (it is unhashable) while a tuple can.
  • Using a set instead of a list inside a loop to cut an O(n²) job down to O(n).
  • Knowing that dictionaries keep insertion order since Python 3.7.

3. Explain the GIL, multithreading, and multiprocessing

Sample answer

Comparison of threads and processes across workload type, memory, GIL limits, and startup cost.

The Global Interpreter Lock (GIL) is a mutex in CPython that lets only one thread run Python bytecode at a time. Threads therefore do not speed up CPU-heavy work. They do help with I/O-bound work, because a thread releases the GIL while it waits on a network call, a disk read, or a database query.

For CPU-bound work, switch to multiprocessing. Each process gets its own interpreter and memory, so processes run in true parallel across cores. The cost is slower startup and the need to pickle data between processes.

Use threads when your code waits, and use processes when your code computes.

Code example

This snippet uses a thread pool for API calls and a process pool for heavy parsing.

import requests
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor

def fetch(url):            # I/O-bound
    return requests.get(url, timeout=10).json()

def parse(chunk):          # CPU-bound
    return [transform(r) for r in chunk]

with ThreadPoolExecutor(max_workers=16) as pool:
    pages = list(pool.map(fetch, urls))

with ProcessPoolExecutor() as pool:
    results = list(pool.map(parse, chunks))

Fetching 500 API pages one by one is slow because the program idles on the network. Sixteen threads overlap those waits. Parsing is different, so the process pool uses every core instead.

What interviewers look for

This question shows up in most python data engineer interview questions because it tests whether you can match a tool to a workload. Listen for these points:

  • A clear split between I/O-bound and CPU-bound tasks, with an example of each from a pipeline.
  • Awareness that asyncio is a third option for many concurrent I/O calls.
  • Knowing that NumPy and pandas release the GIL in many native operations, and that Spark sidesteps the problem by distributing work.

A red flag is the claim that "Python cannot do parallelism." It can. A good answer explains how.

4. Describe decorators and how they help in pipelines

Sample answer

Decorators are functions that wrap another function to add behavior without editing its body. You apply one with the @ syntax above a function definition. In pipelines, they handle cross-cutting concerns such as timing, logging, retries, and caching, so each task function stays focused on its own transformation.

Under the hood, a decorator takes a function and returns a new one, which relies on a closure. Always add functools.wraps, so the wrapped function keeps its name and docstring. That matters for logs and for tools like Airflow that read function metadata.

A decorator lets you write logging, timing, or retry logic once and reuse it on every pipeline step.

Code example

This decorator logs how long each task takes.

import functools, logging, time

def timed(func):
    @functools.wraps(func)
    def wrapper(*args, **kwargs):
        start = time.perf_counter()
        result = func(*args, **kwargs)
        logging.info("%s took %.2fs", func.__name__, time.perf_counter() - start)
        return result
    return wrapper

@timed
def transform(rows):
    return [clean(r) for r in rows]

Every call to transform now logs its runtime, and you added zero lines to the function body. To accept arguments, such as @retry(times=3), you add one more outer layer that returns the decorator.

What interviewers look for

Expect this in python coding interview questions for data engineer roles, often phrased as "write a retry or timing decorator." Strong candidates write one from memory and tie it to a real pipeline need. Listen for these points:

  • Correct use of *args and **kwargs, so the wrapper works with any signature.
  • Using functools.wraps without being prompted.
  • Knowing built-in examples such as @lru_cache, @staticmethod, and @property.

A red flag is a candidate who can define decorators but cannot name a practical use in production.

5. Explain context managers and the with statement

Sample answer

A context manager guarantees that setup and cleanup code run as a pair. The with statement calls __enter__ when the block starts and __exit__ when it ends, even if an exception is raised. That makes it the right tool for files, database connections, locks, and temporary directories.

In pipelines, the payoff is no leaked resources. A forgotten close() on a database connection can exhaust a connection pool after a few thousand runs. You can build your own manager with a class, or more simply with the @contextmanager decorator from contextlib.

A context manager makes cleanup automatic, so a failed task cannot leave a connection or file open.

Code example

This helper wraps a database transaction. It commits on success and rolls back on any error.

from contextlib import contextmanager

@contextmanager
def db_transaction(conn):
    cur = conn.cursor()
    try:
        yield cur
        conn.commit()
    except Exception:
        conn.rollback()
        raise
    finally:
        cur.close()

with db_transaction(conn) as cur:
    cur.execute("INSERT INTO sales VALUES (%s, %s)", (1, 99.5))

Code before yield is the setup. Code after it is the teardown, and the finally block runs no matter what happens inside the with body.

What interviewers look for

Data engineering python interview questions on this topic are short, but they reveal whether you write production-safe code. Listen for these points:

  • An explanation that __exit__ still runs on errors, and that returning True from it suppresses the exception.
  • A generator-based manager written from memory, as in the example.
  • Practical uses beyond files, such as cursors, locks, and temporary directories.

Candidates who only mention open() have shallow experience. Strong ones connect the idea to transactions and cleanup after failed loads.

6. Handle exceptions and retries in ETL scripts

Sample answer

Catch specific exceptions, never a bare except:. Then split failures into two kinds. Transient errors, like a network timeout or a rate limit, deserve a retry. Permanent errors, like a bad schema or a failed login, should fail fast.

Retries need exponential backoff with jitter, so many workers do not hit a struggling service in sync. Cap the attempts, log every failure, and re-raise the last error so the scheduler marks the task as failed.

Retry transient failures, fail fast on permanent ones, and never swallow an error silently.

Code example

Here is a retry decorator that builds on the pattern from question 4.

import functools, random, time

def retry(times=3, base=1.0, exceptions=(ConnectionError, TimeoutError)):
    def decorator(func):
        @functools.wraps(func)
        def wrapper(*args, **kwargs):
            for attempt in range(1, times + 1):
                try:
                    return func(*args, **kwargs)
                except exceptions:
                    if attempt == times:
                        raise
                    time.sleep(base * 2 ** (attempt - 1) + random.random())
        return wrapper
    return decorator

@retry(times=4)
def pull_orders():
    ...

Only ConnectionError and TimeoutError trigger a retry. A ValueError from bad data propagates immediately, and the final attempt re-raises the original error.

What interviewers look for

This is a staple of data engineer python interview questions, because every real pipeline calls flaky systems. Listen for these points:

  • A clear line between transient and permanent errors.
  • A mention of idempotency, since a retried write must not duplicate rows.
  • Bad records routed to a dead-letter file instead of crashing the whole batch.

A red flag is except Exception: pass. It hides data loss, and strong candidates know a silent failure is worse than a loud one.

7. Explain shallow copy vs deep copy

Sample answer

Two folders on a desk, one tied by string to the original envelope, one with a full copy.

A shallow copy builds a new outer object but reuses references to the nested objects inside it. A deep copy recursively duplicates everything, so the copy shares nothing with the original.

Mutable nested data is where the bugs hide. Edit a nested list through a shallow copy and the original changes too. In pipelines, this happens when you copy a config dict for each task and one task's edit leaks into the others. Deep copies cost more time and memory, so use them only when you need full isolation.

A shallow copy duplicates the container, and a deep copy duplicates the container and everything inside it.

Code example

This snippet shows the shared-reference trap.

import copy

config = {"name": "orders", "tables": ["a", "b"]}
shallow = copy.copy(config)
deep = copy.deepcopy(config)

config["tables"].append("c")

print(shallow["tables"])  # ['a', 'b', 'c']
print(deep["tables"])     # ['a', 'b']

The shallow copy still points to the same tables list, so the append shows up in both. The deep copy stays untouched because it owns its own list.

What interviewers look for

Among python interview questions for data engineer roles, this one separates people who memorized definitions from people who have debugged a mutation bug. Listen for these points:

  • Knowing that list.copy(), dict.copy(), and slicing like rows[:] are all shallow.
  • A real example, such as a shared default config or a list of dicts modified inside a loop.
  • Awareness that deepcopy is slow on large objects, and that avoiding mutation is often the better fix.

A red flag is the belief that new = old makes a copy. It only creates a second name for the same object.

8. Read a file larger than memory with pandas

Sample answer

Use the chunksize argument in read_csv. Pandas then returns an iterator of DataFrames instead of one giant frame, so you process the file piece by piece. Aggregate or write out each chunk, then let it go.

You can also cut memory before chunking. Pass usecols to load only the columns you need, and set dtype to shrink column types, such as float32 or category. Past a few tens of GB, a columnar format like Parquet, or a tool like Dask or Spark, is the better choice.

Never load what you can stream: read in chunks, keep only what you need, and write results as you go.

Code example

This snippet totals sales by country from a huge CSV.

import pandas as pd

totals = {}
reader = pd.read_csv(
    "events.csv",
    usecols=["country", "amount"],
    dtype={"country": "category", "amount": "float32"},
    chunksize=500_000,
)
for chunk in reader:
    part = chunk.groupby("country", observed=True)["amount"].sum()
    for country, amt in part.items():
        totals[country] = totals.get(country, 0) + amt

Each chunk holds 500,000 rows. Its partial sums merge into a small dictionary, so peak memory is one chunk plus the totals, not the whole file.

What interviewers look for

This is one of the most common python data engineer coding questions, because it mirrors daily work. Strong candidates name the memory problem first and then pick a fix. Listen for these points:

  • Merging partial results across chunks, since a per-chunk groupby alone gives wrong totals.
  • Using usecols and dtype before reaching for heavier tools.
  • Knowing when to move to Parquet, Dask, or Spark.

A red flag is the answer "just buy more RAM." It shows the candidate has never hit a real memory limit.

9. Clean data and handle missing values in pandas

Sample answer

Start by profiling. Run df.info() and df.isna().sum() to measure missing values before you change anything. Then choose a strategy per column. Use dropna() when few rows are affected, fillna() with a constant, median, or forward fill when the gap is harmless, or add a flag column to record what you imputed.

Cleaning goes beyond nulls. You also standardize types and text, strip whitespace, parse dates, and remove duplicates. Document every rule, because silent imputation can skew downstream metrics.

Decide how to treat each missing value from its business meaning, not from whichever pandas method is shortest.

Code example

This snippet cleans a customer file in a few readable steps.

import pandas as pd

df = pd.read_csv("customers.csv")
df["email"] = df["email"].str.strip().str.lower()
df["signup_date"] = pd.to_datetime(df["signup_date"], errors="coerce")
df["age"] = pd.to_numeric(df["age"], errors="coerce")
df["age_missing"] = df["age"].isna()
df["age"] = df["age"].fillna(df["age"].median())
df["country"] = df["country"].fillna("unknown")
df = df.dropna(subset=["customer_id"]).drop_duplicates("customer_id")

Setting errors="coerce" turns bad values into NaN instead of crashing, so you handle them in one place. Rows without a customer_id are dropped because they can never be joined. The age_missing column preserves a record of which ages were imputed.

What interviewers look for

Questions like this are staples of data engineering python coding questions, because every raw source is dirty. Strong candidates ask why the data is missing before picking a fix. Listen for these points:

  • The difference between NaN, None, and an empty string.
  • Choosing median over mean when outliers exist.
  • Comparing row counts before and after cleaning, and logging what was dropped.

A red flag is filling every null with 0. It corrupts averages and counts, and a good candidate says so.

10. Merge, join, and group data with pandas

Sample answer

Use pd.merge() to combine tables on keys. The how argument sets the join type: inner keeps matches only, left keeps every row from the left table, and outer keeps everything. Add validate="many_to_one" so pandas raises an error on unexpected duplicate keys instead of silently multiplying your rows.

Grouping follows split, apply, combine. You call groupby(), then use agg() with named aggregations to produce several metrics in one pass. Use transform() when you need the result aligned back to the original rows, such as a per-group average beside each record.

Check row counts before and after every join, because a bad key quietly duplicates data.

Code example

This snippet joins orders to customers and reports revenue by country.

import pandas as pd

orders = pd.read_csv("orders.csv")
customers = pd.read_csv("customers.csv")

merged = orders.merge(
    customers[["customer_id", "country"]],
    on="customer_id",
    how="left",
    validate="many_to_one",
)

summary = (
    merged.groupby("country", as_index=False)
    .agg(revenue=("amount", "sum"), orders=("order_id", "nunique"))
)

A left join keeps orders whose customer is missing, so you can count unmatched rows with merged["country"].isna().sum(). The validate flag guards against fan-out from duplicate customer IDs.

What interviewers look for

Python data engineering interview questions on pandas often hide a trap: a join that inflates row counts. Strong candidates test key uniqueness first. Listen for these points:

  • Knowing when to use inner, left, right, and outer joins.
  • Comparing row counts before and after a merge.
  • Telling merge, index-based join, and concat apart.
  • Choosing agg() or transform() based on the output shape.

Candidates who never mention duplicate keys have probably not shipped a join into production.

11. Write an idempotent pipeline

Sample answer

An idempotent pipeline produces the same result whether it runs once or five times for the same input. Reruns and retries are normal in production, so a job that appends duplicates on its second run is broken by design.

Two patterns cover most cases. Delete and insert by partition replaces one day's data inside a single transaction. An upsert relies on a unique key, such as INSERT ... ON CONFLICT DO UPDATE in PostgreSQL. Also keep runs deterministic. Pass the run date in as a parameter, and never call datetime.now() inside the load logic.

Design every load so that running it twice leaves the table exactly as running it once would.

Code example

Here is a daily load that reuses the transaction helper from question 5.

def load_day(conn, run_date, rows):
    with db_transaction(conn) as cur:
        cur.execute("DELETE FROM sales WHERE sale_date = %s", (run_date,))
        cur.executemany(
            "INSERT INTO sales (order_id, sale_date, amount) VALUES (%s, %s, %s)",
            [(r["order_id"], run_date, r["amount"]) for r in rows],
        )

Rerun it for the same date and the old rows are removed first, so no duplicates appear. Because the delete and insert share one transaction, a crash leaves the previous data intact.

What interviewers look for

Python data engineer interview questions on pipelines often end here, because the topic ties together retries, transactions, and scheduling. Strong candidates raise idempotency before you ask. Listen for these points:

  • Choosing between partition overwrite and upsert, and explaining why.
  • Safe backfills, where rerunning last month changes nothing unexpectedly.
  • Staging tables or atomic file renames, so readers never see half-written output.

A red flag is a plain INSERT with no key or partition. That candidate has probably never survived a backfill.

12. Solve common coding challenges: duplicates and lookups

Sample answer

Most python coding questions for data engineer interviews reduce to two moves: dedupe with a set or dict, and look up with a dict. Say your approach out loud before you type. To remove duplicates while keeping order, track seen keys in a set. To keep the latest record per key, load rows into a dictionary so later rows overwrite earlier ones.

Lookups follow the same logic. Instead of scanning a list for every item, build a dictionary once and query it. That turns an O(n²) nested loop into a single O(n) pass.

Reach for a set or dict first, because hash lookups turn nested loops into single passes.

Code example

This snippet solves three classics: ordered deduplication, the newest row per ID, and finding two numbers that hit a target sum.

def dedupe(items):
    seen, out = set(), []
    for x in items:
        if x not in seen:
            seen.add(x)
            out.append(x)
    return out

def latest_by_id(rows):
    ordered = sorted(rows, key=lambda r: r["updated_at"])
    return list({r["id"]: r for r in ordered}.values())

def two_sum(nums, target):
    index = {}
    for i, n in enumerate(nums):
        if target - n in index:
            return index[target - n], i
        index[n] = i

Sorting by updated_at first means the dictionary keeps the newest row for each ID. In two_sum, each number checks the dictionary for its complement, so the list is read only once.

What interviewers look for

These problems are short, so interviewers judge how you think, not how fast you type. Listen for these points:

  • Asking whether order matters and which duplicate to keep before writing code.
  • Stating time and space complexity unprompted.
  • Testing edge cases such as empty input, null keys, and repeated values.
  • Mentioning pandas drop_duplicates or SQL ROW_NUMBER() as the production equivalent.

A red flag is nested loops with in list checks on large inputs. It shows the candidate has never watched a job crawl on real data.

Preparing for your interview

These 12 questions share one thread: memory, failure, and repeatability. Whether the topic is generators, pandas joins, or an idempotent load, interviewers want to hear that you think about scale and about what happens when something breaks. Work through the python interview questions for data engineer roles above by typing each solution from scratch. Then explain it out loud in under two minutes. That habit beats memorizing answers.

Recruiters can flip the same list into a structured screening round. Ask three or four of these questions, score the answers against the "what interviewers look for" points, and you will quickly see who has shipped real pipelines. If you are hiring, search 180,000+ verified candidate profiles on Olibr and shortlist data engineers by skill, experience, and location.

For engineers

Find work worth your time.

Live engineering roles across India and the US, matched to your stack. Build a profile recruiters actually discover.

Browse jobsHow it works

O
§ The author

Olibr Team

Reviewed by Raman Gupta, Founder, Olibr

Filed underHiring Tips
Reading time20 min · 3,815 words

PublishedOctober 5, 2026

CategoryHiring Tips
Enjoyed this piece?Share it with someone who would find it useful.
§ Stay in the loop

Don’t miss the next one.

We publish essays on engineering, hiring, and building teams. Subscribe and we’ll send them when they land.

Unsubscribe anytime · one letter, never more