Problem Set 1

Problem Set Due Wednesday September 23rd at 7pm on Canvas.

You will hand in a .Rmd file and a knitted html output.

I have provided the raw RMD of this problem set you can use as a template for answering.

📝 Download the RMD template for this problem set

Question 1: Probability & Counting

  1. In a game of poker each player is dealt 5 cards from a 52 card deck. How many different 5 card poker hands can be generated from a 52 card deck? (Hint: Is this a permutation or a combination?).

  2. There are 4 suits in a card deck, each consisting of 13 cards. A “flush” is a poker hand where all 5 cards are of the same suit. Calculate the probability of being dealt a flush. (Hint: You first need to count all the ways you can make a 5 card hand using only cards from one suit.)

  3. Simulate the answer to question (b) using R. Run a loop 100000 times that draws 5 cards from a deck of 52, where there are 4 suits with 13 cards each. Note, that we don’t care which cards are which within a suit. We just need 13 hearts, 13 diamonds, 13 clubs, 13 spades. How often are you dealt a flush? (Hint: to see if all the cards in my hand were of the same suit I started with the unique() command.)

  4. A manufacturer of code-based locks comes to you worried that the codes on his lock are too easy to guess. He tells you that his locks have a dial with 10 numbers and the codes are 3 digits long. Like most locks, the numbers in the code cannot repeat and the order of the numbers matters. Calculate the probability of guessing this code using both math and simulation. For the simulation, run the loop 1 million times.

  5. To help this manufacturer we want to determine if it’s more effective to manufacture a bigger dial or to require the user to use a longer code. Using R to simulate each possibility 1 million times, make two graphs. For the first graph, determine the probability of guessing a code with dial sizes from 10 to 30 numbers and a 3 digit code. For the second graph, determine the probability of guessing a code with a dial size of 10, but code lengths from 3-10 numbers long. What do you find?

Question 2: Conditional Probability

The dataset apps is (fake) data on 100,000 college applicants. We have indicator (dummy) variables for high.sat (whether they scored high on the SAT or not), admit (whether they were admitted to the college or not), and high.performer (whether they will perform highly in college or not).

apps <- rio::import("https://github.com/marctrussler/IIS-Data/raw/refs/heads/main/PS1CondProb.Rds", trust=T)
  1. What share of all applicants are high performers?

  2. How does achieving a high SAT score affect performance? Use the conditional probability formula (\(P(A|B) = \frac{P(A\&B)}{P(B)}\)) to calculate \(P(HighPerform|HighSAT)\) and \(P(HighPerform|LowSAT)\)

  3. Now calculate \(P(HighPerformer|HighSAT \& Admitted)\) and \(P(HighPerformer|LowSAT \& Admitted)\), that is the conditional probability of being a high performer based on SAT scores among only the admitted students. What do you find?

  4. Can you explain the paradoxical result you just found? Try to think it through and answer, though if you want a hint you can look up Berkson’s Paradox.

Question 3: Properties of a random variable

Consider the following probability mass function of a random variable, \(K\):

\[ f(K) = p(K=k) = \begin{cases} \frac{1}{6} \text{ if } 1\\ \frac{1}{3} \text{ if } 2\\ \frac{1}{3} \text{ if } 4\\ \frac{1}{6} \text{ if } 10\\ \end{cases} \]

  1. What is the CDF of \(k\)?

  2. What is the expected value of K?

  3. What is the variance of K?

  4. Use R to draw 10,000 samples of K. Confirm that the expected value and variance you calculated above is roughly correct.

  5. Plot the PMF and CDF of K, comparing the simulated and calculated values. For the CDF try out the cumsum() function.