Course Outline
Day 01
Strategic Overview of Big Data Business Intelligence for Criminal Intelligence Analysis
- Criminal justice case studies: Applications in predictive policing
- Adoption trends of Big Data within law enforcement and alignment of operational strategies with predictive analytics
- Emerging technological solutions, including gunshot detection systems, video surveillance, and social media monitoring
- Leveraging Big Data technologies to manage and mitigate information overload
- Integrating Big Data architectures with legacy data systems
Fundamentals of enabling technologies in predictive analytics
Data Integration & Dashboard Visualization
- Strategies for fraud management
- Application of business rules in fraud detection
- Threat detection and suspect profiling techniques
- Cost-benefit analysis frameworks for Big Data implementation for government agencies
Introduction to Big Data
- Core characteristics of Big Data: Volume, Variety, Velocity, and Veracity
- Massively Parallel Processing (MPP) architecture
- Data Warehouses: Static schemas and slowly evolving datasets
- MPP Database systems: Greenplum, Exadata, Teradata, Netezza, Vertica, and others
- Hadoop-based solutions: Handling datasets without predefined structural constraints
- Standard operational pattern: HDFS storage, MapReduce processing, and data retrieval from HDFS
- Apache Spark for real-time stream processing
- Batch processing: Optimized for analytical and non-interactive workloads
- Volume: Complex Event Processing (CEP) for streaming data
- Industry-standard CEP products: Infostreams, Apama, MarkLogic, among others
- Emerging technologies with limited production readiness: Storm/S4
- NoSQL Databases (columnar and key-value): Complementary analytical tools to traditional data warehouses
NoSQL Solutions
- Key-Value Stores: Keyspace, Flare, SchemaFree, RAMCloud, Oracle NoSQL Database
- Key-Value Stores: Dynamo, Voldemort, Dynomite, SubRecord, MongoDB, DovetailDB
- Hierarchical Key-Value Stores: GT.m, Cache
- Ordered Key-Value Stores: TokyoTyrant, Lightcloud, NMDB, Luxio, MemcacheDB, Actord
- Key-Value Caching: Memcached, Repcached, Coherence, Infinispan, EXtremeScale, JBossCache, Velocity, Terracotta
- Tuple Stores: Gigaspaces, Coord, Apache River
- Object Databases: ZopeDB, DB40, Shoal
- Document Stores: CouchDB, Cloudant, Couchbase, MongoDB, Jackrabbit, XML-Databases, ThruDB, CloudKit, Persistent, Riak-Basho, Scalaris
- Wide Columnar Stores: BigTable, HBase, Apache Cassandra, Hypertable, KAI, OpenNeptune, Qbase, KDI
Data Variety and Data Cleaning Challenges in Big Data
- RDBMS limitations: Static structures/schemas that hinder agile, exploratory environments
- NoSQL advantages: Semi-structured data handling, allowing storage without pre-defined schemas
- Key challenges in data cleaning processes
Hadoop Ecosystem
- Criteria for selecting Hadoop as a solution
- Structured Data: Enterprise data warehouses can store massive volumes at high cost but enforce rigid structures, limiting active exploration
- Semi-structured Data: Challenging to process with traditional DW/DB solutions
- Data Warehousing: Requires significant effort and results in static models post-implementation
- Big Data Volume and Variety: Processed on commodity hardware using Hadoop
- Hadoop Cluster Requirements: Commodity hardware infrastructure
Introduction to MapReduce and HDFS
- MapReduce: Distributing computational tasks across multiple servers
- HDFS: Ensuring local data availability for computing processes with redundancy
Data characteristics: Unstructured or schema-less formats, differing from traditional RDBMS
- Developer responsibility for interpreting and contextualizing raw data
- MapReduce programming involves Java development considerations and manual data loading into HDFS
Day 02
Big Data Ecosystem: Constructing Big Data ETL (Extract, Transform, Load) -- Selecting Appropriate Tools
- Comparative analysis: Hadoop versus other NoSQL solutions
- Ideal applications for interactive, random data access
HBase: A column-oriented database built on top of Hadoop
- Support for random data access with specific constraints (maximum 1 PB)
- Optimization: Suitable for logging, counting, and time-series data; less ideal for ad-hoc analytics
- Sqoop: Facilitating import from databases to Hive or HDFS via JDBC/ODBC
- Flume: Streaming data, such as log files, into HDFS
Big Data Management System Components
- ZooKeeper: Managing configuration, coordination, and naming services for dynamic compute nodes
- Oozie: Orchestrating complex pipelines, workflows, dependencies, and sequential tasks
- Ambari: Facilitating deployment, configuration, cluster management, and upgrades (system administration)
Cloud-based Management: Whirr
Predictive Analytics: Fundamental Techniques and Machine Learning-Based Business Intelligence
- Introduction to Machine Learning principles
- Classification techniques for machine learning
- Bayesian Prediction methods for training file preparation
- Support Vector Machines (SVM)
- KNN p-Tree Algebra and vertical mining
Neural Networks applications
- Addressing the large variable problem in Big Data: Random Forests (RF)
- Solving Big Data automation challenges with Multi-model Ensemble RF
Automation strategies via Soft10-M
- Text analysis utilizing Treeminer tools
Agile learning methodologies
- Agent-based learning models
- Distributed learning frameworks
- Overview of open-source predictive analytics tools: R, Python, RapidMiner, Mahout
Predictive Analytics Ecosystem and Applications in Criminal Intelligence Analysis for government operations
- Integration of technology into the investigative process
Insight analytic methodologies
- Visualization analytics
- Structured predictive analytics
- Unstructured predictive analytics
- Profiling of threats, fraudsters, and vendors
- Recommendation engine functionalities
- Pattern detection mechanisms
- Rule and scenario discovery: Identifying failures, fraud, and optimization opportunities
- Root cause analysis
- Sentiment analysis techniques
CRM analytics applications
- Network analytics
Text analytics for extracting insights from transcripts, witness statements, and online communications
- Technology-Assisted Review (TAR) processes
- Fraud analytics strategies
- Real-time analytic capabilities
Day 03
Real-Time and Scalable Analytics Over Hadoop Infrastructure
- Analysis of why standard analytic algorithms fail in Hadoop/HDFS environments
- Apache Hama: Facilitating Bulk Synchronous distributed computing
Apache Spark: Enabling cluster computing and real-time analytics
- CMU Graphics Lab2: Graph-based asynchronous approaches to distributed computing
- KNN p: Algebraic approaches from Treeminer for reducing hardware operational costs
Tools for eDiscovery and Digital Forensics for government legal proceedings
- Comparative analysis of cost and performance: eDiscovery on Big Data vs. Legacy systems
Predictive coding and Technology Assisted Review (TAR)
- Live demonstration of vMiner to illustrate how TAR accelerates discovery processes
- Achieving faster indexing through HDFS to enhance data velocity
NLP (Natural Language Processing): Open-source products and techniques
- eDiscovery in foreign languages: Technologies for processing multilingual content
Big Data Business Intelligence for Cyber Security: Achieving a 360-degree view, rapid data collection, and threat identification for government cybersecurity initiatives
- Fundamentals of security analytics: Understanding attack surfaces, security misconfigurations, and host defenses
Network infrastructure and large data pipeline management for real-time response ETL
- Differentiating prescriptive vs. predictive analytics: Moving from fixed rule-based systems to auto-discovery of threat rules via metadata
Gathering Disparate Data Sources for Criminal Intelligence Analysis for government investigations
- Utilizing IoT (Internet of Things) sensors for data acquisition
- Leveraging satellite imagery for domestic surveillance purposes
Using surveillance and image data for criminal identification
- Additional data gathering technologies: Drones, body-worn cameras, GPS tagging systems, and thermal imaging
- Integrating automated data retrieval with information from informants, interrogations, and research
Forecasting criminal activity patterns
Day 04
Fraud Prevention Business Intelligence through Big Data in Fraud Analytics for government enforcement
- Classification of Fraud Analytics: Rules-based versus predictive analytics approaches
Supervised vs. unsupervised machine learning for detecting fraud patterns
- Types of fraud: Business-to-business, medical claims, insurance, tax evasion, and money laundering
Social Media Analytics: Intelligence Gathering and Analysis for government monitoring
- Analyzing criminal use of social media for organization, recruitment, and planning
Big Data ETL APIs for extracting social media data
- Processing text, images, metadata, and video content
Sentiment analysis derived from social media feeds
- Contextual and non-contextual filtering of social media streams
- Social media dashboards for integrating diverse platforms
Automated profiling of social media accounts
- Live demonstrations of analytic capabilities using Treeminer tools
Big Data Analytics in Image Processing and Video Feeds for government situational awareness
- Image storage techniques for Big Data: Solutions for petabyte-scale data
- LTFS (Linear Tape File System) and LTO (Linear Tape Open)
GPFS-LTFS (General Parallel File System - Linear Tape File System): Layered storage for large image datasets
- Fundamentals of image analytics
Object recognition technologies
- Image segmentation techniques
- Motion tracking capabilities
3-D image reconstruction methods
Biometrics, DNA Analysis, and Next Generation Identification Programs for government justice systems
- Advancements beyond fingerprinting and facial recognition
Speech recognition, keystroke dynamics, and CODIS (Combined DNA Index System)
- Beyond standard DNA matching: Forensic DNA phenotyping for constructing facial approximations from samples
Big Data Dashboards for Rapid Accessibility and Visualization of Diverse Data for government decision-making:
- Integration of existing application platforms with Big Data dashboards
Big Data management strategies
- Case studies: Tableau and Pentaho dashboard implementations
Promoting location-based services in government through Big Data applications
- Tracking system and asset management functionalities
Day 05
Justifying Big Data Business Intelligence Implementation Within an Organization:
- Defining Return on Investment (ROI) for Big Data initiatives
Case studies demonstrating time savings for analysts in data collection and preparation, thereby increasing productivity
- Revenue gains from reduced database licensing costs
Revenue opportunities through location-based services
- Cost savings achieved via fraud prevention measures
- An integrated spreadsheet methodology for calculating approximate expenses versus revenue gains/savings from Big Data implementation for government budgeting.
Step-by-Step Procedure for Migrating From Legacy Data Systems to Big Data Systems
- Big Data Migration Roadmap development
Critical information requirements before architecting a Big Data system
- Methodologies for calculating Volume, Velocity, Variety, and Veracity of data
- Estimating future data growth projections
- Relevant case studies
Review of Major Big Data Vendors and Their Product Offerings.
- Accenture
- APTEAN (Formerly CDC Software)
- Cisco Systems
- Cloudera
- Dell
- EMC
- GoodData Corporation
- Guavus
- Hitachi Data Systems
- Hortonworks
- HP
- IBM
- Informatica
- Intel
- Jaspersoft
- Microsoft
- MongoDB (Formerly 10Gen)
- MU Sigma
- Netapp
- Opera Solutions
- Oracle
- Pentaho
- Platfora
- Qliktech
- Quantum
- Rackspace
- Revolution Analytics
- Salesforce
- SAP
- SAS Institute
- Sisense
- Software AG/Terracotta
- Soft10 Automation
- Splunk
- Sqrrl
- Supermicro
- Tableau Software
- Teradata
- Think Big Analytics
- Tidemark Systems
- Treeminer
- VMware (Part of EMC)
Q&A Session
Requirements
- Familiarity with law enforcement procedures and data infrastructure
- Foundational proficiency in SQL/Oracle or relational database management systems
- Competency in statistical analysis at the spreadsheet application level
Target Audience
- Law enforcement professionals possessing technical expertise, tailored for government personnel
Testimonials (3)
basics and loved the prepared documents and exercises
Rekha Nallam - GE Medical Systems Polska Sp. z o.o.
Course - Introduction to Predictive AI
Deepthi was super attuned to my needs, she could tell when to add layers of complexity and when to hold back and take a more structured approach. Deepthi truly worked at my pace and ensured I was able to use the new functions /tools myself by first showing then letting me recreate the items myself which really helped embed the training. I could not be happier with the results of this training and with the level of expertise of Deepthi!
Deepthi - Invest Northern Ireland
Course - IBM Cognos Analytics
he was well prepared - and he is very sympathetic