Senior Data Analyst Graph

DS STREAM REMOTE WARSAW KRAKÓW GDAŃSK WROCŁAW 2026-08-24
The Senior Data Analyst is a hands-on analytical contributor embedded in the engineering
team. This role bridges the gap between the data itself and the engineers, product
managers, and data analysts building and evaluating the Knowledge Graph. The primary
focus is data integrity: validating datasets, verifying query outputs, tracing the root cause of
discrepancies, and applying statistical methods to assess data quality across multiple
storage technologies.
This is a practitioner role, not a consulting engagement. The deliverable is evidence —
validated results, documented defects, root cause analysis, and statistical assessments
that the team can act on.
  • Practical experience applying statistical methods to data quality assessment:
    distribution analysis, outlier detection, variance analysis, sampling validation
  • Ability to interpret benchmark result data and distinguish meaningful performance
    differences from noise
    Cross-Technology Proficiency
  • Comfort working across multiple database technologies and query languages —
    this role will need to query data in PostgreSQL, graph databases, and Databricks as
    part of normal validation work
  • Experience with Databricks or similar distributed data platforms (Spark, Delta Lake)
    Communication and Collaboration
  • Strong written communication — validation findings, defect reports, and root cause
    analyses must be clear enough for both engineers and product stakeholders
  • Ability to work independently under minimal supervision, taking direction from
    peers rather than requiring structured management oversight
  • Experience embedded in a cross-functional engineering team
  • Strongly Preferred Qualifications
  • Hands-on experience with graph databases (Neo4j, TigerGraph, or similar) — even
    at proof-of-concept scale
  • Familiarity with graph data models: property graphs, node/edge schema,
    relationship taxonomies
  • Domain knowledge in compliance, KYC/AML, or financial services data — beneficial
    ownership structures, sanctions screening, PEP designation, corporate ownership
    chains
  • Familiarity with knowledge graph or ontology concepts
,[Validate datasets loaded into each technology to confirm completeness, accuracy, and structural integrity relative to source data in Databricks and CDS , Verify benchmark query outputs across technologies — confirm that the same logical query against the same underlying data produces consistent, correct results regardless of which system executes it , Identify, document, and trace the root cause of data discrepancies and defects discovered during validation; distinguish between ETL issues, schema translation errors, technology-specific behavior, and upstream data quality problems , Develop and maintain validation test cases and expected outputs for benchmark queries and compliance use cases Use Case Assessment Support , Support the collection and documentation of compliance use cases from the business unit , Help assess each use case: can it be fulfilled by a conventional database approach, or does it require the traversal and pattern-matching capabilities of a purpose-built graph database , Contribute analytical rigor to use case triage — this is a cost-and-complexity decision as much as a technical one ] Requirements: TigerGraph, Neo4j Tools: . Additionally: Sport subscription, Private healthcare, Flat structure, Small teams, International projects, Modern office, Startup atmosphere, No dress code.