All jobs
B

Junior Data Engineer (São Paulo, BR)

Blab
Brazil August 17, 2026
StandardsCertification & Product Delivery
Applying to this role?

Tailor your resume to this exact posting and check it against the ATS — free.

Key skills & keywords for this role

These are the terms most likely to matter to the ATS for this Junior Data Engineer (São Paulo, BR) role at Blab. Mirror the ones that match your real experience on your resume to improve your match score.

StandardsCertification & Product DeliveryquotfontDataspanstylesize12ptfamilyhelveticaarial

About the role

<div class=&quot;content-intro&quot;><hr> <p>&nbsp;</p></div><p><em><span style=&quot;font-weight: 400;&quot;>This is a <strong>Full-Time</strong> Role (40 hours per week, 5 days per week) with no option for part-time work. While this is a remote-first opportunity, the candidate filling this role must be a resident of Pennsylvania, New York, or Brazil at the start of employment. Additionally, they must be within commuting distance of our office in Philadelphia, New York City, or São Paulo.&nbsp;<br></span></em><em>Please visit <a href=&quot;https://boards.greenhouse.io/blab&quot; target=&quot;_blank&quot;>our Careers page</a> to review all opportunities and submit your application for the role(s) that best fit your location and work authorization.</em></p> <hr> <h2><strong>About the Team</strong></h2> <p><span style=&quot;font-size: 12pt; font-family: helvetica, arial, sans-serif;&quot;>The Data &amp; AI team is responsible for building B Lab s data platform, delivering on business priorities, and scaling AI adoption across the organization. The team reports directly to the CTDO. The Data &amp; ML Platforms team owns the data platform, ML infrastructure, and the foundational data models the rest of the Data &amp; AI team depends on — the supply-side capability layer that Business &amp; Data Priorities and AI Enablement build on.</span></p> <h2><strong>About the Opportunity</strong></h2> <p><span style=&quot;font-size: 12pt;&quot;>As a Junior <strong>Data Engineer</strong> within the Data &amp; ML Platforms pillar, you build and maintain the data pipelines and infrastructure the rest of the Data &amp; AI team depends on.</span><br><span style=&quot;font-size: 12pt;&quot;>Your first priority is unblocking the onboarding of new data sources — currently one of the team s top hiring priorities, paused pending this hire. You will absorb data engineering work currently split between the Senior Machine Learning Engineer and the Pillar Lead, freeing them to focus on ML infrastructure and platform strategy respectively.</span><br><span style=&quot;font-size: 12pt;&quot;>You will work closely with the Senior Analytics Engineer on core data modeling and with the Data Governance Lead on data standards, definitions, and compliance.</span></p> <h2><strong>Core Responsibilities&nbsp;</strong></h2> <p><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Data Pipeline Development &amp; New Source Onboarding (55%):</span></p> <ul> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Own Pipeline Development End-to-End: Design, build, and maintain robust, scalable ETL/ELT pipelines that reliably deliver clean data to the platform.</span></li> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Add New Data Sources: Evaluate, scope, and integrate new data sources as they re identified.</span></li> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Partner Across Pillars: Work continuously with Network Priorities, Regional Enablement and AI Enablement to understand incoming data needs and translate them into pipeline work.</span></li> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Monitor &amp; Troubleshoot: Proactively identify and resolve pipeline failures and data quality issues before they affect downstream users.</span></li> </ul> <p><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Platform Support &amp; Cross-Team Collaboration (35%):</span></p> <ul> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Support Foundational Data Models: Work with the Senior Analytics Engineer to maintain the core data models the rest of the team depends on.</span></li> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Ensure Data Availability for Consumers: Make sure the data needed by Data Analysts and the Senior Machine Learning Engineer is reliably available.</span></li> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Follow Data Governance Standards: Apply the data standards, definitions, and sensitivity classifications set by the Data Governance Lead.</span></li> </ul> <p><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Strategic Innovation &amp; Business Impact (10%):</span></p> <ul> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Evaluate Pipeline Tooling: Explore and pilot new ETL/ELT tools or approaches that could improve onboarding speed or pipeline reliability.</span></li> <li style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;><span style=&quot;font-family: helvetica, arial, sans-serif; font-size: 12pt;&quot;>Quantify Impact: Track and articulate how pipeline reliability and onboarding speed affect downstream analytics and ML work.</span><br><br></li> </ul> <h3>Major Objectives/Project for the role in the first 6-12 months</h3> <ul> <li>Add new data sources to the Data Platform</li> <li>Improve data infrastructure allowing for less downtime and more proactive monitoring</li> <li>Enable transformation of data and data models&nbsp;</li> <li>Improve dependency handling between ta
Apply on arbeitnow Posting aggregated by WeZoom · applications happen on the source site.