joshuavogel.dev

Available for select engagements — Q4 2026

Your data platform is slow, expensive, and distrusted. I fix that.

I'm Joshua Vogel — a data engineer and web developer who helps growth-stage companies cut warehouse spend, rebuild pipeline reliability, and ship semantic layers their executives actually trust. I can also help you POC or deploy your app. Fixed-scope engagements. Quantified outcomes.

years in data
12+
years in data
avg. cost reduction
38%
avg. cost reduction
platforms shipped
40+
platforms shipped

Resume

Filter by technology

Languages & Scripting

Bash

BI & Data Platforms

AI/ML & Architecture

RESTful APIs

Additional Tools & Domains

  1. Business impact

    90%

    Faster report release cycles

    Dashboard deployment automation replaced manual release work

    200%

    Increase in model generation

    Self-serve reporting removed release bottlenecks

    3x

    Faster model builds

    Nextflow orchestration improved build performance

    Technical achievements

    • Replaced hours-long manual Airtable workflows with a Python-based ontology management and reporting system that standardized metadata, preserved historical accuracy, and ensured consistent data availability for ML model generation.
    • Accelerated new report release cycles by 90% by building a Tableau deployment automation tool that used the Tableau API to clone and republish dashboards across BigQuery datasets with repeatable configuration.
    • Created a self-serve reporting framework that removed BI bottlenecks and drove a 200% increase in model generation throughput across the company.
    • Improved model build performance by 3x by scripting data processing and pipeline orchestration in Nextflow, making builds more predictable, resumable, and easier to scale.
    • Clarified the company's overarching data model by combining hive partitioning with granular data improving data discoverability and query performance.
    • Improved data trust by enforcing strict dbt modeling discipline, source testing, and relationship validation, which eliminated critical-level support tickets tied to broken or inconsistent data contracts.

    Responsibilities

    • Administered BigQuery datasets, access patterns, partitioning strategy, and reporting-layer reliability for the ML platform.
    • Served as the Tableau administrator, managing dashboard deployments, access controls, and automation workflows.
    • Maintained the accuracy of the model's metadata via update scripts and automated validation checks.
  2. Business impact

    >75%

    Faster reporting loads

    SQL architecture and modeling changes reduced dashboard latency

    10%

    Cloud spend reduction

    Usage monitoring and query alerts reduced waste

    15%

    Faster product delivery

    Executive dashboards improved resource visibility and planning

    Technical achievements

    • Re-engineered backend SQL in BigQuery with more efficient CTE patterns, materialized views, and optimized data models to reduce reporting load times by more than 75% for internal stakeholders.
    • Designed a real-time operational dashboard suite that surfaced high-cost queries and heavy usage trends, helping reduce Google Cloud spend by 10% company-wide.
    • Built a Jira-to-BigQuery ETL pipeline in TypeScript to ingest project and resource data for executive dashboards, improving delivery speed by 15% and giving leadership better portfolio visibility.
    • Served as the central owner for Tableau dashboard architecture, optimization, and stakeholder communication across departments, enabling better decision-making at the executive level.

    Responsibilities

    • Owned the BI platform and translated executive and cross-functional reporting needs into maintainable data products.
  3. Business impact

    50%

    Faster record matching

    Improved matching throughput for reimbursement workflows

    5 hours

    Weekly manual work removed

    Automated exception routing eliminated repetitive list generation

    80%

    Manual workload reduction

    Automation reduced operating burden for call-center and clinical teams

    Technical achievements

    • Increased record-matching speeds by 50% and laid the groundwork for a predictive pricing AI model by implementing an NLP-based entity resolution workflow for insurance plan matching.
    • Converted brittle CSV-based workflows into SQLite-backed data stores, enabling faster querying, anomaly detection, and operational diagnostics for reimbursement trend analysis while preventing file corruption issues.
    • Cut five hours of weekly manual list generation by building an automated routing pipeline in Python that sent unmatched data exceptions directly to the clinical call team.
    • Used data quality monitoring and exception handling to improve the reliability of pricing and reimbursement datasets used by upstream operational teams.
  4. Business impact

    >20%

    Increase in physician goal adherence

    Automated tracking and operational monitoring improved compliance

    1/5

    Lower 30-day readmission rate

    Risk-based intervention model targeted high-cost cohorts

    Technical achievements

    • Built an automated performance tracking system to monitor provider adherence on key operational metrics such as patient leakage and preventive care, improving goal adherence by more than 20%.
    • Used targeted SQL models to identify high-risk, high-cost patient cohorts for post-hospitalization follow-up, helping reduce 30-day readmission rates by one-fifth.
    • Translated operational needs into reliable, repeatable reporting and analysis that supported risk and performance management decisions across the business.

    Responsibilities

    • Owned recurring risk and performance reporting for provider adherence, patient leakage, preventive care, and follow-up operations.
    • Translated business questions into SQL analyses and cohort definitions for clinical and operational stakeholders.
    • Maintained reliable reporting workflows used to monitor goals and prioritize interventions across high-cost patient populations.
  5. Business impact

    10+

    Languages documented

    Extensive documentation of Jewish languages

    5000+

    Monthly active users

    Active collaboration with communities for language preservation

    Technical achievements

    • Established normative CI/CD pipelines to streamline development and deployment processes.
    • Edited the Lexicon codebase to improve readability and maintainability as well as compliance with CakePHP conventions.
    • Integrated automated testing to ensure code quality and reliability.
    • Optimized the performance of the Lexicon codebase to handle large datasets efficiently.

    Responsibilities

    • Maintain the network of Jewish Language lexicons across 10+ languages.
    • Implement enhancements to the Lexicon codebase to improve functionality and usability.
    • Pull data from various Jewish language lexicons for research and analysis purposes.

Proof of Work

Featured case studies

All case studies →

Services

Defined scope. Tailored proposal. Measurable outcome.

Full service details →

Typical engagement: 2 weeks

Warehouse & Pipeline Cost Audit

Two-week forensic pass over your warehouse spend. You get a prioritized optimization roadmap — and the quick wins implemented before I leave.

Typical engagement: 4 weeks

dbt & Semantic Layer Design

Four weeks to a metric layer your execs trust: restructured dbt project, certified definitions, and PR-based change management.

Typical engagement: 2–3 weeks

Pipeline Reliability Review

SLIs/SLOs for every pipeline, idempotency audit, and an alerting strategy that pages for real incidents only.

Typical engagement: 3–6 weeks

Web App Build & Deployment

From working prototype to production: I build the app, connect the systems it needs, and set up repeatable deployments your team can own.

One bad quarter of warehouse spend pays for this twice.

Thirty minutes, no pitch deck. Bring your ugliest data problem — I'll tell you honestly whether I can fix it and what it would take.