BT-393 – Data Engineer
Location- Chantilly
**MUST HAVE A TS/SCI CLEARANCE TO APPLY. Those without an active security clearance will not be considered.**
Bespoke Technologies is seeking a Software Developer to provide ETL, Data Engineering, and Full-Stack Software Development support
Required Skills:
- Demonstrates experience designing and maintaining enterprise-grade ETL/ELT pipelines, both batch and real-time.
- Demonstrates front-end development and implementation skills, using React, Next.js, or similar.
- Demonstrates back-end development using Python, Java, Scala, and microservices architecture.
- Demonstrates experience with API design.
- Demonstrates experience with containerization, using Docker.
- Demonstrates experience with CI/CD pipelines.
- Demonstrates experience with infrastructure-as-code patterns.
- Demonstrates experience with probabilistic models, risk scoring, Bayesian inference, Monte Carlo simulation, and probabilistic graphical models.
- Demonstrates experience applying statistical modeling tools.
- Demonstrates experience with designing cloud-native architectures using cloud services such as AWS, Google, IBM, and Oracle
- Demonstrates experience designing and operating big data systems
- Demonstrates experience building and optimizing performance of large scale graph databases (tens of billions of edges) using DynamoDB or new enhanced capabilities
- Demonstrates experience developing and operating graph traversal capabilities using data graphing tool traversal capabilities built upon Apache Gremlin or new enhanced capabilities
- Demonstrates experience developing and operating NoSQL solutions to complex big data applications
- Demonstrates experience in data modeling for performance, partition sharding, record/event aggregation workflows, stream processing, and metrics gathering
- Demonstrates experience designing and operating large-scale serverless geospatial indexes built with GeoMESA
- Demonstrates experience with partition and sort key design and implementation to ensure consistent performance
- Demonstrates experience with aggregation operations to de-duplicate records on continuous data feeds
- Demonstrates subject matter expertise experience with relational databases to noSQL
- Demonstrates experience building and operating high performance data processing pipelines using Lambda, Step Functions and PySpark
- Demonstrates experience building high quality User Interface/User experiences with the React framework and webGL
- Demonstrates experience designing and operating large scale graph databases using Apache Cassandra
- Demonstrates experience performing in-depth technical analysis of large-scale graph databases to develop implementation strategies for search optimizations
- Demonstrates experience developing technical capabilities for processing, persistence and search of datasets that are collected or maintained using standards common in the Sponsor's community
- Demonstrates experience facilitating engineering discussions across teams representing multiple stakeholders to develop and execute implementation strategies that meet mission needs
- Demonstrates experience developing Machine Learning Operations (MLOps) pipelines for large scale application
- Demonstrates experience maintaining configuration of software using configuration management resources such as GitHub
- Demonstrates experience designing, building and operating big data systems, such as persistence, partitioning, indexing, at scale of trillions of records/events
- Demonstrates experience with Niagara Files (NiFi) applications or new enhanced capabilities
- Demonstrates experience developing and operating Kubernetes infrastructure
- Demonstrates experience supporting engineering efforts that will contribute to delivery of capabilities such as datasets and functionality such as communications, geospatial workflows
- Demonstrates experience implementing DevSecOps and agile development in production environments
- Demonstrates experience with agile software development and testing
- Demonstrates experience with federal security, regulatory and compliance requirements and security accreditation package development
- Demonstrates experience with data security and governance using centralized security controls like LDAP, encrypting the data, and auditing access to the data
- Demonstrates experience with specialized technologies that are optimized for the particular use of the data, such as relational databases, a NoSQL database (Cassandra), or object storage
- Demonstrates experience with Apache, TINKERPOP, GREMLIN and/or JANUSGRAPH to design, develop, implement and maintain system
- Demonstrates knowledge of Graph Database to design, develop, implement and maintain system
- Demonstrates experience with C or C++ to write interfaces
- Demonstrates experience using centralized security controls like LDAP, encrypting data, and auditing access to data
- Demonstrates experience with:
- Databases: Postgres, MariaDB, ELK, Minio, AWS S3, Neo4j, MongoDB, noSQL
- Languages: Python (pypi libraries)
- Operating Systems: Centos7, RockyLinux8
- Orchestration: Kubernetes, Docker, Docker-Compose, Docker-Swarm
- Development Tools: vscode, gitlab, jupyterhub/notebooks, MATLAB
- Environments: large collaboration and development environments
- Data types: Unstructured, structured, or semi-structured data, including: CSV, JSON, JSONL, AVRO, Protocol Buffers, Parquet, etc
Desired Skills:
- Demonstrates experience with designing cloud-native architectures using Sponsors cloud services
- Demonstrates experience designing and operating big data systems within the Sponsors policy and regulatory environment
- Demonstrates experience developing and operating graph traversal capabilities using the Sponsors data graphing tool traversal capabilities built upon Apache Gremlin
- Demonstrates experience building and operating high performance data processing pipelines using Lambda, Step Functions and PySpark on the Sponsors infrastructure with EMR
- Demonstrates experience working with the Sponsor's enterprise services used for Data Management, including the enterprise catalog service (and associated APIs), and Policy Decision Points (PDPs).
- Demonstrates experience developing Machine Learning Operations (MLOps) pipelines for large scale application in the Sponsor's environment
- Demonstrates experience and understanding of IT Service Management and common SLA measurements
- Demonstrates experience presenting solutions, requirements, and presentations to diverse audiences.
- Demonstrates experience working with container orchestration technologies such as AWS ECS, AWS Fargate, and Kubernetes or other enhanced capabilities available
- Demonstrates experience in managing large operational cloud environments spanning multiple tenants using Multi-Account management, AWS Well Architected Best Practices, and AWS Organization Units/Service Control Policies (OU/SCP).
- Demonstrates experience with Micro-services such as building decoupled systems, utilizing RESTful endpoints and lightweight systems
- Demonstrates experience in total systems perspectives, including a technical understanding of systems and applications relationships, dependencies, and requirements of hardware and software components
- Demonstrates experience consulting with customers to determine present and future user needs
- Demonstrates experience providing frequent contact with customers, traceability within program documents, and the overall computing environment and architecture
Desired Certifications:
- AWS Certified Solutions Architect
- AWS Machine Learning Certification(s)
- Agile certification
- Azure
- Security+
- GSEC
- CCNA