Major structure

Overview

This major is available in the Bachelor of Science.

In the Data Science major, you’ll build statistical and computational foundations to work with real-world data at scale.

Major structure

The Data Science major is made up of eight subjects (100 credit points) taken in your second and third year. Each subject is worth 12.5 credit points. Level 2 subjects are usually taken in second year, and Level 3 subjects in third year.

To complete this major, you’ll need:

The rest of your degree will consist of a Level 1 science core subject, your choice of elective subjects in science, and breadth (non-science) subjects.

Further information

You can find detailed information about your major – including structure, subject availability, and participation requirements – in the Handbook and on the School of Mathematics and Statistics website.

You can also explore your study pathway and sample course plans through My Course Planner.

Sample course plan

View some sample course plans to help you select subjects that will meet the requirements for this major.

Data Science Start-year intake

These sample study plans assume that students have achieved a study score of at least 29 in VCE Specialist Mathematics 3/4, or equivalent. If students have not completed this previously, they may first need to enrol in MAST10005 Calculus 1 in their first semester.

Year 1

100 pts

Semester 1 · 50 pts
  • Today's Science, Tomorrow's World – core – SCIE10005 – 12.5 pts
  • Calculus 2 – elective – MAST10006 – 12.5 pts
  • Foundations of Algorithms – elective – COMP10002 – 12.5 pts
  • breadth – 12.5 pts
Semester 2 · 50 pts
  • Linear Algebra – elective – MAST10007 – 12.5 pts
  • elective – 12.5 pts
  • elective – 12.5 pts
  • breadth – 12.5 pts

Year 2

100 pts

Semester 1 · 50 pts
  • Elements of Data Processing – major – COMP20008 – 12.5 pts
  • Database Systems – major – INFO20003 – 12.5 pts
  • Probability – major – MAST20004 – 12.5 pts
  • breadth – 12.5 pts
Semester 2 · 50 pts
  • Statistics – major – MAST20005 – 12.5 pts
  • elective – 12.5 pts
  • elective – 12.5 pts
  • breadth – 12.5 pts

Year 3

100 pts

Semester 1 · 50 pts
  • Linear Statistical Models – major – MAST30025 – 12.5 pts
  • Machine Learning – major – COMP30027 – 12.5 pts
  • elective – 12.5 pts
  • breadth – 12.5 pts
Semester 2 · 50 pts
  • Modern Applied Statistics – major – MAST30027 – 12.5 pts
  • Applied Data Science – major – MAST30034 – 12.5 pts
  • elective – 12.5 pts
  • elective – 12.5 pts

Explore this major

Explore the subjects you could choose as part of this major.

Core

Students must complete all required core subjects

Accordion
Elements of Data Processing · 12.5 pts

AIMS

Data processing is fundamental to computing and data science. This subject covers various aspects of data processing including database management, representation and analysis of data, information retrieval, visualisation and reporting, and cloud computing. This subject includes an emphasis on both tools and underlying foundations.

INDICATIVE CONTENT

The subject's focus is on the data pipeline, and activities known colloquially as 'data wrangling'. Indicative topics covered include:

  • Capturing data (data ingress)
  • Data representation and storage
  • Cleaning, normalisation and filling in missing data (imputation)
  • Combing multiple sources of data (data integration)
  • Query languages and processing
  • Scripting to support the data pipeline
  • Visualisation and presentation

View detailed information in the Handbook

Machine Learning · 12.5 pts

AIMS

Machine Learning, a core discipline in data science, is prevalent across Science, Technology, the Social Sciences, and Medicine; it drives many of the products we use daily such as banner ad selection, email spam filtering, and social media newsfeeds. Machine Learning is concerned with making accurate, computationally efficient, interpretable and robust inferences from data. Originally borne out of Artificial Intelligence, Machine Learning has historically been the first to explore more complex prediction models and to emphasise computation, while in the past two decades Machine Learning has grown closer to Statistics gaining firm theoretical footing.

This subject aims to introduce undergraduate students to the intellectual foundations of machine learning, and to introduce practical skills in data analysis that can be applied in graduates' professional careers.

CONTENT

Topics will be selected from: prediction approaches for classification/regression such as k-nearest neighbour, naïve Bayes, discriminative linear models, decision trees, Support Vector Machines, Neural Networks; clustering methods such as k-means, hierarchical clustering; probabilistic approaches; exposure to large-scale learning.

View detailed information in the Handbook

Database Systems · 12.5 pts

AIMS

Contemporary online services such as social networking and multimedia-sharing sites, massive multiplayer online games and commerce services have database management systems at their back-end. In this subject, students will obtain a deep understanding of the concepts behind database management systems. In particular, the students will become familiar with the database system architecture, and will exercise the concepts such as query processing and optimisation, database tuning and transactions, which are the foundation of any modern data processing application. This subject is core within the Bachelor of Science for the Major of Computing and Software Systems and the Major of Informatics. Students completing the Diploma of Informatics are also required to undertake this subject.

INDICATIVE CONTENT

This subject serves as an introduction to data modelling and databases from a technical and data management perspective. The subject will include Entity Relationship modelling (from conceptual design to physical modelling), normalisation, de-normalisation, relational model and relational algebra, SQL, query processing and query optimisation, transactions, storage organisation, database administration, data warehousing and big data analytics. Other topics in data management and DBMS technology with an overview of modern NoSQL systems may also be included.

View detailed information in the Handbook

Statistics · 12.5 pts

This subject introduces the basic elements of statistical modelling, computation and data analysis. It is an entry point to further study of both mathematical and applied statistics, as well as broader data science.

Students will develop the ability to fit statistical models to data, estimate parameters of interest and test hypotheses. Both classical and Bayesian approaches will be covered. The importance of the underlying mathematical theory of statistics and the use of modern statistical software will be emphasised.

Concepts covered include: descriptive statistics, random sample, statistical inference, point estimation, interval estimation, properties of estimators, maximum likelihood, confidence intervals, hypothesis testing and Bayesian inference. Applications covered include: exploratory data analysis, inference for samples from univariate distributions, simple linear regression, correlation, goodness-of-fit tests and analysis of variance.

View detailed information in the Handbook

Linear Statistical Models · 12.5 pts

Linear models are central to the theory and practice of modern statistics. They are used to model a response as a linear combination of explanatory variables and are the most widely used statistical models in practice. Starting with examples from a range of application areas this subject develops an elegant unified theory that includes the estimation of model parameters, quadratic forms, hypothesis testing using analysis of variance, model selection, diagnostics on model assumptions, and prediction. Both full rank models and models that are not of full rank are considered. The theory is illustrated using common models and experimental designs.

View detailed information in the Handbook

Modern Applied Statistics · 12.5 pts

Modern applied statistics combines the power of modern computing and theoretical statistics. This subject considers the computational techniques required for the practical implementation of statistical theory, and includes Bayes and Monte-Carlo methods. The subject focuses on the application of these techniques to generalised linear models, which are commonly used in the analysis of categorical data.

View detailed information in the Handbook

Applied Data Science · 12.5 pts

AIMS

This capstone subject for the Data Science major combines statistical reasoning and practical computing skills to solve challenging problems with big data.

INDICATIVE CONTENT

Students will learn about communication of quantitative information and insights; presentation skills; report writing; project management; problem formulation using case studies; data collection and measurement protocols; data from surveys and experiments; issues in capturing and dealing with “big data”; dimension reduction; data visualisation; fitting formulated models to data to infer insightful information about populations; ethics in quantitative research; working effectively in teams.

View detailed information in the Handbook

Elective

Students select subjects according to the major requirements

Accordion
Probability · 12.5 pts

This subject offers a thorough grounding in the basic concepts of mathematical probability and probabilistic modelling. Topics covered include random experiments and sample spaces, probability axioms and theorems, discrete and continuous random variables/distributions (including measures of location, spread and shape), expectations and generating functions, independence of random variables and measures of dependence (covariance and correlation), methods for deriving the distributions of transformations of random variables or approximations for them (including the central limit theorem).

The probability distributions and models discussed in the subject arise frequently in real world applications. These include a number of widely used one- and two-dimensional (particularly the bivariate normal) distributions and also fundamental probability models such as Poisson processes and Markov chains.

View detailed information in the Handbook

Probability for Statistics · 12.5 pts

This subject develops the probability theory that is necessary to understand statistical inference. Properties of probability are reviewed, random variables are introduced, and their properties are developed and illustrated through common univariate probability models. Models for the joint behavior of random variables are introduced, along with conditional probability and Markov chains. Methods for obtaining the distributions of functions of random variables are considered along with techniques to obtain the exact and approximate distributions of sums of random variables. These methods will be illustrated through some well known normal approximations to discrete distributions and by obtaining the exact and approximate distributions of some commonly used statistics. Computer packages are used for numerical calculations but no programming skills are required.

View detailed information in the Handbook