Get in Touch

Course Outline

Day 01

Strategic Overview of Big Data Business Intelligence for Criminal Intelligence Analysis

  • Criminal justice case studies: Applications in predictive policing
  • Adoption trends of Big Data within law enforcement and alignment of operational strategies with predictive analytics
  • Emerging technological solutions, including gunshot detection systems, video surveillance, and social media monitoring
  • Leveraging Big Data technologies to manage and mitigate information overload
  • Integrating Big Data architectures with legacy data systems
  • Fundamentals of enabling technologies in predictive analytics

Data Integration & Dashboard Visualization

  • Strategies for fraud management
  • Application of business rules in fraud detection
  • Threat detection and suspect profiling techniques
  • Cost-benefit analysis frameworks for Big Data implementation for government agencies

Introduction to Big Data

  • Core characteristics of Big Data: Volume, Variety, Velocity, and Veracity
  • Massively Parallel Processing (MPP) architecture
  • Data Warehouses: Static schemas and slowly evolving datasets
  • MPP Database systems: Greenplum, Exadata, Teradata, Netezza, Vertica, and others
  • Hadoop-based solutions: Handling datasets without predefined structural constraints
  • Standard operational pattern: HDFS storage, MapReduce processing, and data retrieval from HDFS
  • Apache Spark for real-time stream processing
  • Batch processing: Optimized for analytical and non-interactive workloads
  • Volume: Complex Event Processing (CEP) for streaming data
  • Industry-standard CEP products: Infostreams, Apama, MarkLogic, among others
  • Emerging technologies with limited production readiness: Storm/S4
  • NoSQL Databases (columnar and key-value): Complementary analytical tools to traditional data warehouses

NoSQL Solutions

  • Key-Value Stores: Keyspace, Flare, SchemaFree, RAMCloud, Oracle NoSQL Database
  • Key-Value Stores: Dynamo, Voldemort, Dynomite, SubRecord, MongoDB, DovetailDB
  • Hierarchical Key-Value Stores: GT.m, Cache
  • Ordered Key-Value Stores: TokyoTyrant, Lightcloud, NMDB, Luxio, MemcacheDB, Actord
  • Key-Value Caching: Memcached, Repcached, Coherence, Infinispan, EXtremeScale, JBossCache, Velocity, Terracotta
  • Tuple Stores: Gigaspaces, Coord, Apache River
  • Object Databases: ZopeDB, DB40, Shoal
  • Document Stores: CouchDB, Cloudant, Couchbase, MongoDB, Jackrabbit, XML-Databases, ThruDB, CloudKit, Persistent, Riak-Basho, Scalaris
  • Wide Columnar Stores: BigTable, HBase, Apache Cassandra, Hypertable, KAI, OpenNeptune, Qbase, KDI

Data Variety and Data Cleaning Challenges in Big Data

  • RDBMS limitations: Static structures/schemas that hinder agile, exploratory environments
  • NoSQL advantages: Semi-structured data handling, allowing storage without pre-defined schemas
  • Key challenges in data cleaning processes

Hadoop Ecosystem

  • Criteria for selecting Hadoop as a solution
  • Structured Data: Enterprise data warehouses can store massive volumes at high cost but enforce rigid structures, limiting active exploration
  • Semi-structured Data: Challenging to process with traditional DW/DB solutions
  • Data Warehousing: Requires significant effort and results in static models post-implementation
  • Big Data Volume and Variety: Processed on commodity hardware using Hadoop
  • Hadoop Cluster Requirements: Commodity hardware infrastructure

Introduction to MapReduce and HDFS

  • MapReduce: Distributing computational tasks across multiple servers
  • HDFS: Ensuring local data availability for computing processes with redundancy
  • Data characteristics: Unstructured or schema-less formats, differing from traditional RDBMS

  • Developer responsibility for interpreting and contextualizing raw data
  • MapReduce programming involves Java development considerations and manual data loading into HDFS

Day 02

Big Data Ecosystem: Constructing Big Data ETL (Extract, Transform, Load) -- Selecting Appropriate Tools

  • Comparative analysis: Hadoop versus other NoSQL solutions
  • Ideal applications for interactive, random data access
  • HBase: A column-oriented database built on top of Hadoop

  • Support for random data access with specific constraints (maximum 1 PB)
  • Optimization: Suitable for logging, counting, and time-series data; less ideal for ad-hoc analytics
  • Sqoop: Facilitating import from databases to Hive or HDFS via JDBC/ODBC
  • Flume: Streaming data, such as log files, into HDFS

Big Data Management System Components

  • ZooKeeper: Managing configuration, coordination, and naming services for dynamic compute nodes
  • Oozie: Orchestrating complex pipelines, workflows, dependencies, and sequential tasks
  • Ambari: Facilitating deployment, configuration, cluster management, and upgrades (system administration)
  • Cloud-based Management: Whirr

Predictive Analytics: Fundamental Techniques and Machine Learning-Based Business Intelligence

  • Introduction to Machine Learning principles
  • Classification techniques for machine learning
  • Bayesian Prediction methods for training file preparation
  • Support Vector Machines (SVM)
  • KNN p-Tree Algebra and vertical mining
  • Neural Networks applications

  • Addressing the large variable problem in Big Data: Random Forests (RF)
  • Solving Big Data automation challenges with Multi-model Ensemble RF
  • Automation strategies via Soft10-M

  • Text analysis utilizing Treeminer tools
  • Agile learning methodologies

  • Agent-based learning models
  • Distributed learning frameworks
  • Overview of open-source predictive analytics tools: R, Python, RapidMiner, Mahout

Predictive Analytics Ecosystem and Applications in Criminal Intelligence Analysis for government operations

  • Integration of technology into the investigative process
  • Insight analytic methodologies

  • Visualization analytics
  • Structured predictive analytics
  • Unstructured predictive analytics
  • Profiling of threats, fraudsters, and vendors
  • Recommendation engine functionalities
  • Pattern detection mechanisms
  • Rule and scenario discovery: Identifying failures, fraud, and optimization opportunities
  • Root cause analysis
  • Sentiment analysis techniques
  • CRM analytics applications

  • Network analytics
  • Text analytics for extracting insights from transcripts, witness statements, and online communications

  • Technology-Assisted Review (TAR) processes
  • Fraud analytics strategies
  • Real-time analytic capabilities

Day 03

Real-Time and Scalable Analytics Over Hadoop Infrastructure

  • Analysis of why standard analytic algorithms fail in Hadoop/HDFS environments
  • Apache Hama: Facilitating Bulk Synchronous distributed computing
  • Apache Spark: Enabling cluster computing and real-time analytics

  • CMU Graphics Lab2: Graph-based asynchronous approaches to distributed computing
  • KNN p: Algebraic approaches from Treeminer for reducing hardware operational costs

Tools for eDiscovery and Digital Forensics for government legal proceedings

  • Comparative analysis of cost and performance: eDiscovery on Big Data vs. Legacy systems
  • Predictive coding and Technology Assisted Review (TAR)

  • Live demonstration of vMiner to illustrate how TAR accelerates discovery processes
  • Achieving faster indexing through HDFS to enhance data velocity
  • NLP (Natural Language Processing): Open-source products and techniques

  • eDiscovery in foreign languages: Technologies for processing multilingual content

Big Data Business Intelligence for Cyber Security: Achieving a 360-degree view, rapid data collection, and threat identification for government cybersecurity initiatives

  • Fundamentals of security analytics: Understanding attack surfaces, security misconfigurations, and host defenses
  • Network infrastructure and large data pipeline management for real-time response ETL

  • Differentiating prescriptive vs. predictive analytics: Moving from fixed rule-based systems to auto-discovery of threat rules via metadata

Gathering Disparate Data Sources for Criminal Intelligence Analysis for government investigations

  • Utilizing IoT (Internet of Things) sensors for data acquisition
  • Leveraging satellite imagery for domestic surveillance purposes
  • Using surveillance and image data for criminal identification

  • Additional data gathering technologies: Drones, body-worn cameras, GPS tagging systems, and thermal imaging
  • Integrating automated data retrieval with information from informants, interrogations, and research
  • Forecasting criminal activity patterns

Day 04

Fraud Prevention Business Intelligence through Big Data in Fraud Analytics for government enforcement

  • Classification of Fraud Analytics: Rules-based versus predictive analytics approaches
  • Supervised vs. unsupervised machine learning for detecting fraud patterns

  • Types of fraud: Business-to-business, medical claims, insurance, tax evasion, and money laundering

Social Media Analytics: Intelligence Gathering and Analysis for government monitoring

  • Analyzing criminal use of social media for organization, recruitment, and planning
  • Big Data ETL APIs for extracting social media data

  • Processing text, images, metadata, and video content
  • Sentiment analysis derived from social media feeds

  • Contextual and non-contextual filtering of social media streams
  • Social media dashboards for integrating diverse platforms
  • Automated profiling of social media accounts

  • Live demonstrations of analytic capabilities using Treeminer tools

Big Data Analytics in Image Processing and Video Feeds for government situational awareness

  • Image storage techniques for Big Data: Solutions for petabyte-scale data
  • LTFS (Linear Tape File System) and LTO (Linear Tape Open)
  • GPFS-LTFS (General Parallel File System - Linear Tape File System): Layered storage for large image datasets

  • Fundamentals of image analytics
  • Object recognition technologies

  • Image segmentation techniques
  • Motion tracking capabilities
  • 3-D image reconstruction methods

Biometrics, DNA Analysis, and Next Generation Identification Programs for government justice systems

  • Advancements beyond fingerprinting and facial recognition
  • Speech recognition, keystroke dynamics, and CODIS (Combined DNA Index System)

  • Beyond standard DNA matching: Forensic DNA phenotyping for constructing facial approximations from samples

Big Data Dashboards for Rapid Accessibility and Visualization of Diverse Data for government decision-making:

  • Integration of existing application platforms with Big Data dashboards
  • Big Data management strategies

  • Case studies: Tableau and Pentaho dashboard implementations
  • Promoting location-based services in government through Big Data applications

  • Tracking system and asset management functionalities

Day 05

Justifying Big Data Business Intelligence Implementation Within an Organization:

  • Defining Return on Investment (ROI) for Big Data initiatives
  • Case studies demonstrating time savings for analysts in data collection and preparation, thereby increasing productivity

  • Revenue gains from reduced database licensing costs
  • Revenue opportunities through location-based services

  • Cost savings achieved via fraud prevention measures
  • An integrated spreadsheet methodology for calculating approximate expenses versus revenue gains/savings from Big Data implementation for government budgeting.

Step-by-Step Procedure for Migrating From Legacy Data Systems to Big Data Systems

  • Big Data Migration Roadmap development
  • Critical information requirements before architecting a Big Data system

  • Methodologies for calculating Volume, Velocity, Variety, and Veracity of data
  • Estimating future data growth projections
  • Relevant case studies

Review of Major Big Data Vendors and Their Product Offerings.

  • Accenture
  • APTEAN (Formerly CDC Software)
  • Cisco Systems
  • Cloudera
  • Dell
  • EMC
  • GoodData Corporation
  • Guavus
  • Hitachi Data Systems
  • Hortonworks
  • HP
  • IBM
  • Informatica
  • Intel
  • Jaspersoft
  • Microsoft
  • MongoDB (Formerly 10Gen)
  • MU Sigma
  • Netapp
  • Opera Solutions
  • Oracle
  • Pentaho
  • Platfora
  • Qliktech
  • Quantum
  • Rackspace
  • Revolution Analytics
  • Salesforce
  • SAP
  • SAS Institute
  • Sisense
  • Software AG/Terracotta
  • Soft10 Automation
  • Splunk
  • Sqrrl
  • Supermicro
  • Tableau Software
  • Teradata
  • Think Big Analytics
  • Tidemark Systems
  • Treeminer
  • VMware (Part of EMC)

Q&A Session

Requirements

  • Familiarity with law enforcement procedures and data infrastructure
  • Foundational proficiency in SQL/Oracle or relational database management systems
  • Competency in statistical analysis at the spreadsheet application level

Target Audience

  • Law enforcement professionals possessing technical expertise, tailored for government personnel
 35 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories