Job-Nr. 39262774
DH

Senior Data Scientist - (Content, Consumer) (m/f/d)

Delivery Hero SE vor 2 Wochen
VollzeitHybridSenior
StandortBerlin, Berlin

Wondering what relocating to Berlin is like? In this article, we’ve put together 10 things you should know about moving to Berlin and how Delivery Hero can support you. We believe diversity and inclusion are key to creating not only an exciting product, but also an amazing customer and employee experience. Fostering this starts with hiring - therefore we do not discriminate on the basis of racial identities, religious beliefs, color, national origin, gender identities or expressions, sexual orientations, age, marital or disability statuses, or any other aspect that makes you, you.

Severely disabled applicants with equal qualifications will be given preferential consideration.

Your Mission

  • You'll own the LLM systems behind social proof end-to-end — quality, reliability, and coverage across languages, platforms, and use cases — including the unglamorous parts: drift, miscalibration, silent quality degradation, and data issues in production.
  • You'll set the standard for how this team builds with LLMs. Model and provider selection based on empirical cost, latency, and quality evidence. Prompt strategy that is systematic rather than folkloric. The judgment on when an LLM is the wrong tool.
  • You'll drive the roadmap through problem discovery, finding high-impact gaps, quantifying the business value, and turning them into scoped initiatives — treating cost of inference as a product decision, not only an engineering constraint.
  • You'll define and operationalise meaningful metrics, and build the evaluation infrastructure behind them — offline eval suites, LLM-as-judge frameworks, and annotation processes — so that evaluation reflects true business value rather than misleading proxies, and a prompt change or model swap becomes a one-day decision.
  • You'll take prototypes to production, shaping architecture and data flows with backend and data engineering, and building the feedback loops that let the system keep improving without constant manual intervention.
  • You'll raise the bar beyond your own work, through best practices, mentoring, and a culture of ownership and pragmatism.

NLP & LLM systems depth. You have built LLM-based systems that reached production, preferably on noisy, multilingual, user-generated text. You can explain the architecture, the failure modes you hit, and why you made the calls you made. You have worked across more than one model family and more than one architecture, and you can reason empirically about cost, latency, and quality rather than defaulting to a familiar vendor or pattern. You know when to prompt, when to add context or retrieval, when to fine-tune, when a small task-specific model wins, and when not to use an LLM at all.

  • Evaluation rigour for generative systems. You know an online A/B test is too slow and too blunt to iterate an LLM system on. You have designed offline evaluation suites, used LLM-as-judge patterns while being honest about their biases, structured annotation for generative outputs, and measured faithfulness, hallucination, and quality at scale.
  • Observability for LLM pipelines. You instrument what you build: tracing chained calls, watching token usage and cost, and capturing intermediate outputs so a multi-step pipeline can actually be debugged. You have a view on where this is non-negotiable and where it is overhead.
  • Product thinking and problem framing. You turn an ambiguous business need into a well-scoped DS problem, spot the high-impact opportunity nobody put on the roadmap, and connect your technical work to outcomes people outside data science care about.
  • Ambiguity and an MVP instinct. You move from zero to one on imperfect information, ship in increments, and resist over-engineering.
  • Leverage beyond yourself. You improve how the people around you work — practices, standards, rigour — and you are the reason a team gets better at something, not just faster at it.

Also valuable

  • Production-grade Python and analytical SQL, with familiarity with ML lifecycle practices, orchestration, monitoring, and modern engineering standards such as version control and CI/CD.
  • A strong foundation in statistics and causal inference, and the instinct to challenge a result that looks too good.
  • Experience with data collection and labelling pipelines, including working with annotation teams.
  • Agentic development tools used pragmatically — including applying them to model improvement itself, such as automated prompt optimisation or LLM-assisted annotation — with the judgment to know when to verify the output.

Additional Information

Ensuring you and all our Heroes are looked after, happy, and healthy is always on the menu. Because if you’re in good shape, then we’re in good shape.

  • Make the most of our hybrid working model and join the team for face-to-face connection and collaboration in our beautiful Berlin campus 2 days a week
  • We offer 27 days holiday with an extra day on 2nd and 3rd year of service
  • We will support you in developing yourself and your career growth opportunities: 1.000 € Educational Budget, Language Courses, Parental Support and access to the Udemy Business platform to explore a variety of online courses.
  • Get moving and release those wonderful, mind-boosting endorphins: Health Checkups, Meditation & Gym.
  • Cash. Dough. Cheddar. Whatever you call it, we’ll help you with it: Employee Share Purchase Plan, Sabbatical Bank, Public Transportation Ticket Discount, Life & Accident Insurance, Corporate Pension Plan
  • The power of getting together over some food is unrivaled. Here are a few ways to help you do that. All the yum: Digital Meal Vouchers and Food Vouchers.

Why is this one different? You define the direction. Architecture, model strategy, evaluation approach — these are open questions, and they become yours. You are the LLM authority for the squad, not a contributor to someone else's. Part of the job is raising everyone else's ceiling. Real scale, real consequence. Your models move conversion across Delivery Hero's global platforms. Genuinely unsolved problems. Multilingual UGC at scale, empirical multi-provider model selection, prompt optimisation, and evaluation infrastructure for generative output — none of these have a settled answer here or anywhere. Agentic development is a first-class part of how we work, used pragmatically to move faster without losing system understanding.

You're welcome to share your pronouns (he/she/they) right from the start so we can address you respectfully from our first contact.

Über den Arbeitgeber

DH
Delivery Hero SE
Unternehmen · Berlin
9
offene Stellen
Alle Stellen des Arbeitgebers

Könnte dich auch interessieren