Bespoke Technologies
Posted 20h ago

Data Engineer

Bespoke Technologies
Chantilly, Virginia, United States
OnsiteFull Time
Responsibilities
  • designing pipelines
  • operating databases
  • developing interfaces
Requirements
  • Requires active TS/SCI clearance and experience with ETL
  • Python
  • Graph and NoSQL databases
  • Cloud architectures
  • Big data
  • Kubernetes
  • Docker, CI/CD
  • Security, and data pipelines
Technical tools mentioned
ReactNext.jsPythonJavaScalaDockerAWSGoogleIBMOracleDynamoDBApache GremlinGeoMESADLambdaStep FunctionsPySparkWebGLApache CassandraGitHubNiagara Files (NiFi)KubernetesLDAPApacheTINKERPOPJANUSGRAPHCC++PostgresMariaDBELKMinioAWS S3Neo4jMongoDBnoSQLCentos7RockyLinux8Docker-ComposeDocker-Swarmvscodegitlabjupyterhub/notebooksMATLABCSVJSONJSONLAVROProtocol BuffersParquetAWS ECSAWS FargateEMRRESTful

Job description

BT-393 – Data Engineer
Location- Chantilly



**MUST HAVE A TS/SCI CLEARANCE TO APPLY. Those without an active security clearance will not be considered.**





Bespoke Technologies is seeking a Software Developer to provide ETL, Data Engineering, and Full-Stack Software Development support

Required Skills:

  • Demonstrates experience designing and maintaining enterprise-grade ETL/ELT pipelines, both batch and real-time.
  • Demonstrates front-end development and implementation skills, using React, Next.js, or similar.
  • Demonstrates back-end development using Python, Java, Scala, and microservices architecture.
  • Demonstrates experience with API design.
  • Demonstrates experience with containerization, using Docker.
  • Demonstrates experience with CI/CD pipelines.
  • Demonstrates experience with infrastructure-as-code patterns.
  • Demonstrates experience with probabilistic models, risk scoring, Bayesian inference, Monte Carlo simulation, and probabilistic graphical models.
  • Demonstrates experience applying statistical modeling tools.
  • Demonstrates experience with designing cloud-native architectures using cloud services such as AWS, Google, IBM, and Oracle
  • Demonstrates experience designing and operating big data systems
  • Demonstrates experience building and optimizing performance of large scale graph databases (tens of billions of edges) using DynamoDB or new enhanced capabilities
  • Demonstrates experience developing and operating graph traversal capabilities using data graphing tool traversal capabilities built upon Apache Gremlin or new enhanced capabilities
  • Demonstrates experience developing and operating NoSQL solutions to complex big data applications
  • Demonstrates experience in data modeling for performance, partition sharding, record/event aggregation workflows, stream processing, and metrics gathering
  • Demonstrates experience designing and operating large-scale serverless geospatial indexes built with GeoMESA
  • Demonstrates experience with partition and sort key design and implementation to ensure consistent performance
  • Demonstrates experience with aggregation operations to de-duplicate records on continuous data feeds
  • Demonstrates subject matter expertise experience with relational databases to noSQL
  • Demonstrates experience building and operating high performance data processing pipelines using Lambda, Step Functions and PySpark
  • Demonstrates experience building high quality User Interface/User experiences with the React framework and webGL
  • Demonstrates experience designing and operating large scale graph databases using Apache Cassandra
  • Demonstrates experience performing in-depth technical analysis of large-scale graph databases to develop implementation strategies for search optimizations
  • Demonstrates experience developing technical capabilities for processing, persistence and search of datasets that are collected or maintained using standards common in the Sponsor's community
  • Demonstrates experience facilitating engineering discussions across teams representing multiple stakeholders to develop and execute implementation strategies that meet mission needs
  • Demonstrates experience developing Machine Learning Operations (MLOps) pipelines for large scale application
  • Demonstrates experience maintaining configuration of software using configuration management resources such as GitHub
  • Demonstrates experience designing, building and operating big data systems, such as persistence, partitioning, indexing, at scale of trillions of records/events
  • Demonstrates experience with Niagara Files (NiFi) applications or new enhanced capabilities
  • Demonstrates experience developing and operating Kubernetes infrastructure
  • Demonstrates experience supporting engineering efforts that will contribute to delivery of capabilities such as datasets and functionality such as communications, geospatial workflows
  • Demonstrates experience implementing DevSecOps and agile development in production environments
  • Demonstrates experience with agile software development and testing
  • Demonstrates experience with federal security, regulatory and compliance requirements and security accreditation package development
  • Demonstrates experience with data security and governance using centralized security controls like LDAP, encrypting the data, and auditing access to the data
  • Demonstrates experience with specialized technologies that are optimized for the particular use of the data, such as relational databases, a NoSQL database (Cassandra), or object storage
  • Demonstrates experience with Apache, TINKERPOP, GREMLIN and/or JANUSGRAPH to design, develop, implement and maintain system
  • Demonstrates knowledge of Graph Database to design, develop, implement and maintain system
  • Demonstrates experience with C or C++ to write interfaces
  • Demonstrates experience using centralized security controls like LDAP, encrypting data, and auditing access to data
  • Demonstrates experience with:
  • Databases: Postgres, MariaDB, ELK, Minio, AWS S3, Neo4j, MongoDB, noSQL
  • Languages: Python (pypi libraries)
  • Operating Systems: Centos7, RockyLinux8
  • Orchestration: Kubernetes, Docker, Docker-Compose, Docker-Swarm
  • Development Tools: vscode, gitlab, jupyterhub/notebooks, MATLAB
  • Environments: large collaboration and development environments
  • Data types: Unstructured, structured, or semi-structured data, including: CSV, JSON, JSONL, AVRO, Protocol Buffers, Parquet, etc

Desired Skills:
  • Demonstrates experience with designing cloud-native architectures using Sponsors cloud services
  • Demonstrates experience designing and operating big data systems within the Sponsors policy and regulatory environment
  • Demonstrates experience developing and operating graph traversal capabilities using the Sponsors data graphing tool traversal capabilities built upon Apache Gremlin
  • Demonstrates experience building and operating high performance data processing pipelines using Lambda, Step Functions and PySpark on the Sponsors infrastructure with EMR
  • Demonstrates experience working with the Sponsor's enterprise services used for Data Management, including the enterprise catalog service (and associated APIs), and Policy Decision Points (PDPs).
  • Demonstrates experience developing Machine Learning Operations (MLOps) pipelines for large scale application in the Sponsor's environment
  • Demonstrates experience and understanding of IT Service Management and common SLA measurements
  • Demonstrates experience presenting solutions, requirements, and presentations to diverse audiences.
  • Demonstrates experience working with container orchestration technologies such as AWS ECS, AWS Fargate, and Kubernetes or other enhanced capabilities available
  • Demonstrates experience in managing large operational cloud environments spanning multiple tenants using Multi-Account management, AWS Well Architected Best Practices, and AWS Organization Units/Service Control Policies (OU/SCP).
  • Demonstrates experience with Micro-services such as building decoupled systems, utilizing RESTful endpoints and lightweight systems
  • Demonstrates experience in total systems perspectives, including a technical understanding of systems and applications relationships, dependencies, and requirements of hardware and software components
  • Demonstrates experience consulting with customers to determine present and future user needs
  • Demonstrates experience providing frequent contact with customers, traceability within program documents, and the overall computing environment and architecture
 
Desired Certifications:
  • AWS Certified Solutions Architect
  • AWS Machine Learning Certification(s)
  • Agile certification
  • Azure
  • Security+
  • GSEC
  • CCNA

About Bespoke Technologies

Provides technical engineering and data solutions for national security.

Year founded
2017
Employees
94
Organization type
Private
Latest investment
Funding Round (2017)
Headquarters
US

Similar jobs

Data Engineer roles near Chantilly, Virginia
8h
Save
Mark Applied
Hide
Data Engineer
Arlington, Virginia, United States
$62k-$141k/yr HybridFull Time
Booz Allen Hamilton
Booz Allen HamiltonNYSE: BAH: Consulting and technology services for government and commercial clients
2+ YOEBachelor's degree, Secret clearance, 2+ years with C++, Java, or Python, scalable data stores, and 1+ year with Git or Atlassian tools; ETL/ELT and database experience required.
IoT, C++, Java, Python, Git, Atlassian, SQL, GraphQL, AWS EMR, Redshift, SageMaker, Databricks, SQL Data Warehouse, Apache Spark, NVIDIA CUDA, Terraform, CloudFormation, AWS Solutions Architect, Azure, Security+, CISSP
20h
Save
Mark Applied
Hide
Database Administrator
Windsor Mill, Maryland, United States
$63k-$92k/yr RemoteFull Time
RELI Group
RELI Group: Provides professional IT and management services to government agencies.
Experience designing ETL/ELT pipelines and data lakehouse architectures using Python, SQL, PySpark, AWS, Databricks, and Delta Lake; collaboration, data governance, and CI/CD experience required.
Python, SQL, PySpark, AWS, Databricks, Delta Lake, GitHub, Databricks Repos, CI/CD
3d
Save
Mark Applied
Hide
Senior Data Engineer (AWS, Azure, GCP)
Reston, Virginia, United States
$90k-$200k/yr HybridFull Time
CapTech
CapTech: Provides technology and management consulting services to large enterprises.
5+ YOERequires 5+ years delivering cloud data engineering solutions, expertise in SQL and databases, programming, ETL/orchestration, cloud platforms, data warehousing, distributed systems, and technical leadership.
AWS, Azure, GCP, Azure Data Factory, SSIS, Informatica, Alteryx, Ab Initio, Pentaho, Talend, Matillion, Snowflake, Redshift, Databricks, PostgreSQL, MySQL, SQL Server, Oracle, Aurora, Presto, BigQuery, SQL, NoSQL, Python, Java, R, C, C#, C++, Shell, Git, Jenkins, CI/CD, Jira
3d
Save
Mark Applied
Hide
Data Engineer
Fairfax or United States
$95k-$115k/yr RemoteFull Time, Contract
Bixal
Bixal: Provides digital transformation and consulting services for government agencies.
4+ YOEBachelor’s degree and 4+ years of data engineering experience with Python, PySpark, Databricks, AWS, SQL, Git, Terraform, Linux, data pipelines, and cross-functional engineering; must obtain Public Trust clearance.
PySpark, Databricks Intelligence Platform, AWS, Amazon S3, Apache Kafka, Terraform, Git, Linux, SQL, Apache Hive, Apache Hadoop, Apache Spark, QuickSight, Microsoft Power BI
3d
Save
Mark Applied
Hide
Data Engineer - Axion
Sunnyvale or Washington or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$200k-$275k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
5+ YOERequires 5+ years of relevant experience, modern ML infrastructure, large-scale GPU jobs, data software, microservices or databases, U.S. citizenship, and eligibility for security clearance.
React, TypeScript, Python, Golang, Docker, Kubernetes, Opensearch, Postgres, LLMs, VLMs, GPU
3d
Save
Mark Applied
Hide
Senior Data Engineer - TS/SCI
Bethesda, Maryland, United States
$131k-$237k/yr HybridFull Time
Sunayu
Sunayu: Engineering technology and DevSecOps solutions for defense and intelligence missions.
10+ YOEBachelor's with 12–15 years or master's with 10–13 years of relevant experience; active TS/SCI and polygraph eligibility; expertise in Elasticsearch, graph databases, data pipelines, Kubernetes, Python or Java.
Elasticsearch, OpenSearch, JanusGraph, Neo4j, TigerGraph, Amazon Neptune, Memgraph, Kubernetes, Python, Java, Apache Kafka, Linux, Keycloak, CloudFormation, Terraform, Pulumi, Microsoft Excel
3d
Save
Mark Applied
Hide
Data Engineer
Bethesda, Maryland, United States
OnsiteFull Time
Analytica
Analytica: Provides data analytics and IT solutions for federal government agencies.
5+ YOEBachelor's degree in a quantitative field, 5+ years in data modeling, strategic data planning, governance, and enterprise data warehousing; ETL, SQL Server, communication skills, and ability to obtain Secret clearance required.
Microsoft SQL Server, Amazon Redshift, DBeaver, Microsoft Visual Studio, AWS Glue, AWS Lambda, Git, GitHub, CI/CD, ETL
4d
Save
Mark Applied
Hide
Senior Data Engineer
Ashburn, Virginia, United States
$104k-$173k/yr HybridFull Time
ManTech
ManTech: Provides technology solutions for defense and intelligence agencies.
12+ YOERequires extensive data engineering experience, including data warehousing, relational databases, ETL, cloud migration, communication, and problem-solving. U.S. citizenship and DHS CBP suitability required.
GitLab, Oracle, MySQL, Postgres, SQL Server, Extract-Transform-Load (ETL), Amazon Web Services (AWS), Microsoft Azure, Power BI, OBIEE, Hadoop, HDFS, YARN, Hive, Pig, Spark, Kafka, Storm, Jira, Confluence