Home/Job List/Gen AI Python Developer (Data integration & evaluation engineer)
Crisil

Gen AI Python Developer (Data integration & evaluation engineer)

Crisil

Hyderabad, Telangana, India
Full-Time
Posted 2 days ago

Job Description & Responsibilities

We are seeking a versatile engineer to build the data foundation of our agentic AI analytics platform, built on

LangGraph, Azure OpenAI, AWS Bedrock, Databricks and Postgres and to establish how we measure the quality of

what that platform produces. Our workflows combine market data, proprietary S&P Global datasets, and analystbuilt models into repeatable analytical products used across S&P Global Energy and in client-facing settings. Two

things determine whether those products can be trusted at scale: reliable, structured, well-governed access to the

underlying data, and objective evidence that the output is correct. This role owns both.

You will design and maintain the pipelines, databases, and integrations that allow AI agents to work directly with

trusted energy market data. You will also build the reference datasets, scoring methods, and regression tests that

establish how a change to a prompt, a model, or a retrieval step affects output quality. You will work alongside the

platform and release engineer on the data layer, alongside the agentic AI engineer on evaluation, and alongside

subject matter experts in refining, crude, and related markets to understand what the data means, not only how it is

shaped.

Responsibilities

Data and integration

  • Design, build, and maintain data pipelines that deliver structured, AI-ready data to analytical workflows

and AI agents.

  • Integrate internal and external data sources through APIs, cloud data platforms, and modern agent-tool

protocols, including credential, entitlement, and environment configuration.

  • Design, migrate, and maintain relational database schemas supporting workflow persistence, metadata, and

analytical outputs.

  • Extend and modernize analytical and refining databases as coverage grows, with automated update

processes and clear technical documentation.

  • Normalize and structure raw data so that it can be consumed directly by AI agents within automated

workflows.

  • Build monitoring and validation that maintains data quality across sources and refresh cycles, including

completeness, consistency, and timeliness checks.

  • Work with subject matter experts and analysts to understand domain data, its provenance, and its

limitations.

  • Document data flows, schemas, and integration points so that the work is transferable and auditable.
  • Support deployment and security readiness of data components, including access control, entitlements, and

audit requirements.

Evaluation and quality

  • Design and build evaluation harnesses that score AI-generated output systematically for accuracy,

completeness, consistency, and source traceability.

  • Establish ground-truth and gold-standard reference sets in partnership with subject matter experts,

including training and holdout methodology.

  • Build regression test suites so that changes to prompts, models, retrieval, or workflow logic are validated

ahead of release.

  • Run structured comparisons across AI models and configurations, measuring quality, latency, and compute

cost.

  • Define and detect the characteristic failure modes of generative systems, including unsupported statements,

incomplete field population, conflicting values, and retrieval drift.

  • Build discrepancy detection that surfaces conflicts between sources as explicit, reviewable flags.
  • Work with domain experts to translate expert quality judgments into testable, repeatable criteria.
  • Report findings clearly to both engineers and senior leadership, and drive prioritization of the resulting

improvements.

  • Supply evidence of systematic validation to governance, risk, and security review processes.

S&P Global External

August 2026

Core Requirements

All candidates should be able to demonstrate the following, whatever path they took to acquire it:

  • Proficiency in Python and SQL, sufficient to build both data pipelines and test harnesses and to analyze the

results independently.

  • Practical experience building data pipelines or automated data processes that other people relied on in

production, including what happened when they failed.

  • Working knowledge of relational databases and data modeling, including schema change against systems

with live consumers.

  • Demonstrated experience evaluating, validating, or testing analytical output, models, or systems against a

defined standard.

  • Sound grasp of statistics and experimental design, including sampling, controls, baselines, and sources of

bias.

  • Practical understanding of generative AI applications and prompt engineering, shown through something

you have built rather than tools you have tried.

  • Intellectual honesty and precision: a willingness to report an unwelcome result and defend the method that

produced it.

  • Clear written communication, including the ability to explain a measurement or a data model to someone

who did not design it.

Backgrounds We Will Consider

  • Data or analytics engineering with production ownership of the pipelines you built.
  • AI or machine learning evaluation, large language model evaluation, or applied data science.
  • Model validation, model risk management, or quantitative audit in banking, insurance, energy, or a similar

regulated setting.

  • Scientific or academic research background with strong experimental methodology, in any discipline.
  • Refining, process, or engineering role in which you validated models or simulations against plant or market

reality and write code.

  • Software quality or test automation engineering with genuine analytical depth.
  • Market or research analyst with strong quantitative method, coding ability, and a documented habit of

checking things.

Degrees in engineering, the sciences, statistics, mathematics, economics, computer science, or a related field are all

relevant, as is equivalent practical experience without a matching degree. Method and evidence matter more to us in

this role than any particular credential or industry.

Preferred Skills

  • Postgres, Databricks, or Azure data services run in production.
  • Familiarity with agentic orchestration frameworks such as LangGraph or LangChain, retrieval-augmented

generation, or prompt versioning.

  • Direct experience evaluating large language model or generative AI output, including with tooling such as

LangSmith and Ragas.

  • Experience with model documentation, governance frameworks, or regulatory validation standards.
  • Understanding of token and compute cost economics in AI systems.
  • Experience working in a refinery or chemical industry, or in commodities markets, or with energy domain

data such as crude and refined products, trade flows, or price assessment.

  • Experience in developing and deploying machine learning models in a business context.
  • Experience with data visualization and reporting tools, such as Power BI, Tableau, or Matplotlib

Required Skills

PythonAWSAzureSQLMachine LearningPower BITableauLeadership

Job Details

Employment TypeFull-Time
Work ModeOn-Site
Experience00 years
Positions1

Posted by

N/A

Posted on:

16 Sept 2026

About Crisil

More open roles

Browse all jobs →