FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Hiring · 5 minute read

How to Hire ETL Developers: Signals, Tests and Scope

ETL developers move and reshape data between systems, and the work is judged on whether the numbers are right and stay right. Test for idempotent rerun design, handling of late-arriving and corrected data, and reconciliation habits rather than tool familiarity, since tools change and those problems do not.

By FISTA Solutions· AI-Native Engineering Team·
How to Hire ETL Developers: Signals, Tests and Scope article cover

ETL work is judged on one thing: whether the numbers are right, and whether they stay right when something goes wrong upstream. Tool familiarity is the least durable part of the skill set. This guide covers what to test instead, drawing on FISTA Solutions' staff augmentation and AI enablement work.

What is the first thing to test?

Rerun safety. Ask what happens if a pipeline runs twice on the same day, or is rerun for a date three weeks ago.

PatternRerun behaviour
Append-only insertDuplicates on rerun
Delete-and-reload by partitionSafe, bounded
Merge on natural keySafe if key is genuinely unique
Incremental by timestampMisses late-arriving rows
Full reloadSafe, expensive

Pipelines that duplicate on rerun are extremely common, and the resulting wrong numbers are discovered weeks later by someone who trusted a report.

Why is late-arriving data so hard?

Because it invalidates results already published. A transaction backdated after a period closed, or a correction to a source record, means yesterday's numbers change.

Handling that deliberately — restating, versioning, or flagging — rather than by accident is what distinguishes experienced candidates. Ask what they did when a finance team noticed a figure had changed.

How do you evaluate schema drift handling?

Ask what happens when an upstream system adds, removes, or repurposes a column. Silent failures here are the classic cause of wrong data.

Strong candidates describe contracts, schema checks, and alerting. Weaker ones describe discovering the change when a downstream report broke. See what is data quality.

Why is silent wrongness worse than failure?

Because a failed pipeline gets noticed and fixed within hours, while a pipeline producing subtly wrong numbers gets used for decisions.

Ask what checks they build. Row count variance, null rate changes, referential integrity, and totals reconciled against the source are the practical ones, and candidates who run them routinely produce trustworthy data.

How do you test reconciliation thinking?

Ask how they proved a migrated dataset matched its source. The answer should involve counts, sums on key measures, and spot checks on edge cases — not "we compared a few rows".

Reconciliation is unglamorous and it is the difference between data people trust and data people work around.

What about slowly changing dimensions?

Ask how they handled a record whose attributes change over time when history matters. It is a standard problem with standard approaches, and candidates who have not met it have not built reporting pipelines.

Does tool familiarity matter?

Less than data reasoning. Orchestration and transformation tooling changes every few years; idempotency, late data, drift, and reconciliation are permanent.

Hire for the durable skills and budget a few weeks of tool adjustment.

How does the transformation layer change the role?

Modern practice pushes transformation into the warehouse with version-controlled, tested models, which makes the work closer to software engineering than it used to be.

Ask whether their transformations were tested and reviewed. Candidates from a tooling background where transformations lived in a graphical designer may find that shift substantial.

What about performance and cost?

Warehouse compute is billed, and pipeline design drives most of it. Ask what they did to reduce a pipeline's cost and what they measured.

Candidates who have never considered pipeline cost will build correct and expensive systems.

Contract, staff augmentation, or permanent hire?

Augmentation suits migrations, remediation, and defined pipeline builds. Permanent hiring suits organisations where data is continuously changing and institutional knowledge of what the fields mean is valuable.

What are the common hiring mistakes?

Screening on tool names. Ignoring rerun safety. Omitting reconciliation. And treating pipelines as a delivery project rather than an operated system.

How do you onboard them well?

Give them the pipeline inventory, the freshness dashboard if one exists, and the last three data incidents. Those three describe the estate's real health.

How does AI change this work?

AI systems consume the same pipelines, and they are less forgiving of quiet wrongness than a dashboard is, because a model will confidently use bad data. Data quality checks become more valuable rather than less. See AI enablement.

What does good look like after 90 days?

Critical pipelines safe to rerun, schema checks alerting before downstream breakage, reconciliation running on the datasets that matter, and freshness measured against a stated target.

What should be measured?

Dataset freshness against target, reconciliation variance against source systems, rerun safety coverage, and incidents caused by upstream change.

What should you do first?

Take your three most important pipelines and rerun one for a past date in a test environment. What happens tells you what kind of hire you need.

How do you handle source systems you cannot change?

Most sources are vendor products or legacy applications whose extracts arrive in whatever shape they arrive. Ask what the worst source they worked with did, and how they contained it.

Answers involving a staging layer that lands raw data untouched, with all interpretation applied afterwards, indicate experience. Engineers who transform during extraction lose the ability to reprocess when they discover the interpretation was wrong.

How FISTA Solutions helps

FISTA Solutions builds and staffs data pipelines through staff augmentation and forward deployed engineers: rerun safety treated as a design requirement, late-arriving and corrected data handled deliberately, schema contracts alerting before breakage, reconciliation built in rather than added after a discrepancy, warehouse cost owned as an engineering responsibility, and pipelines built to standards AI systems can consume safely through AI enablement. The record is 150+ projects for 50+ companies across 12+ countries.

To add data pipeline capacity, message FISTA on WhatsApp, or read hire data engineers in Pakistan.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What is the first thing to test?

Rerun safety. Ask what happens if a pipeline runs twice on the same day or is rerun for a past date. Pipelines that duplicate rows or double-count on rerun are extremely common and produce wrong numbers that nobody notices for weeks.

02Why is late-arriving data so hard?

Because it invalidates results already published. A transaction backdated after a period closed, or a correction to a source record, means yesterday's numbers change. Handling that deliberately, rather than by accident, is what distinguishes experienced candidates.

03Does tool familiarity matter?

Less than data reasoning. Orchestration and transformation tools change every few years, while idempotency, late data, schema drift, and reconciliation are permanent. Hire for the second set and expect a few weeks of tool adjustment.

04Why is silent wrongness worse than failure?

Because a failed pipeline gets noticed and fixed, while a pipeline producing subtly wrong numbers gets used for decisions. Strong candidates build checks that fail loudly rather than pipelines that degrade quietly.

05What should be measured?

Dataset freshness against target, reconciliation variance against source systems, and the proportion of pipelines that are safe to rerun. Rows processed measures activity rather than correctness.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project